Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

We introduce methods for discovering and applying sparse feature circuits. These are causally implicated subnetworks of human-interpretable features for explaining language model behaviors. Circuits identified in prior work consist of polysemantic and difficult-to-interpret units like attention heads or neurons, rendering them unsuitable for many downstream applications. In contrast, sparse feature circuits enable detailed understanding of unanticipated mechanisms. Because they are based on fine-grained units, sparse feature circuits are useful for downstream tasks: We introduce SHIFT, where we improve the generalization of a classifier by ablating features that a human judges to be task-irrelevant. Finally, we demonstrate an entirely unsupervised and scalable interpretability pipeline by discovering thousands of sparse feature circuits for automatically discovered model behaviors.

Deep ReinforcementLearning from Human…Deep Reinforcement Learning from Human PreferencesThe Pile: An 800GBDataset of Diverse Text…The Pile: An 800GB Dataset of Diverse Text for Language ModelingLocating and EditingFactual Associations in…Locating and Editing Factual Associations in GPTJumping Ahead: ImprovingReconstruction Fidelity…Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse AutoencodersImproving DictionaryLearning with Gated…Improving Dictionary Learning with Gated Sparse AutoencodersAtP*: An efficient andscalable method for…AtP*: An efficient and scalable method for localizing LLM behaviour to componentsFine-Tuning EnhancesExisting Mechanisms: A…Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity TrackingInterpreting CLIP'sImage Representation vi…Interpreting CLIP's Image Representation via Text-Based DecompositionSuccessor Heads:Recurring, Interpretabl…Successor Heads: Recurring, Interpretable Attention Heads In The WildThe Alignment Problemfrom a Deep Learning…The Alignment Problem from a Deep Learning PerspectiveScaling and evaluatingsparse autoencodersScaling and evaluating sparse autoencodersCausal Abstraction: ATheoretical Foundation…Causal Abstraction: A Theoretical Foundation for Mechanistic InterpretabilityTranscoders FindInterpretable LLM…Transcoders Find Interpretable LLM Feature CircuitsRefusal in LanguageModels Is Mediated by a…Refusal in Language Models Is Mediated by a Single DirectionMeasuring Progress inDictionary Learning for…Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game ModelsA is for Absorption:Studying Feature…A is for Absorption: Studying Feature Splitting and Absorption in Sparse AutoencodersDo I Know This Entity?Knowledge Awareness and…Do I Know This Entity? Knowledge Awareness and Hallucinations in Language ModelsScaling and evaluatingsparse autoencodersScaling and evaluating sparse autoencodersLearning Multi-LevelFeatures with Matryoshk…Learning Multi-Level Features with Matryoshka Sparse AutoencodersTowards PrincipledEvaluations of Sparse…Towards Principled Evaluations of Sparse Autoencoders for Interpretability and ControlNot All Language ModelFeatures Are…Not All Language Model Features Are One-Dimensionally LinearDecomposing The DarkMatter of Sparse…Decomposing The Dark Matter of Sparse AutoencodersInterpreting visiontransformers via…Interpreting vision transformers via residual replacement modelPosition: MechanisticInterpretability Should…Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEsSparse Feature Circuits:Discovering and Editing…Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。