Transcoders Find Interpretable LLM Feature Circuits

A key goal in mechanistic interpretability is circuit analysis: finding sparse subgraphs of models corresponding to specific behaviors or capabilities. However, MLP sublayers make fine-grained circuit analysis on transformer-based language models difficult. In particular, interpretable features -- such as those found by sparse autoencoders (SAEs) -- are typically linear combinations of extremely many neurons, each with its own nonlinearity to account for. Circuit analysis in this setting thus either yields intractably large circuits or fails to disentangle local and global behavior. To address this we explore transcoders, which seek to faithfully approximate a densely activating MLP layer with a wider, sparsely-activating MLP layer. We introduce a novel method for using transcoders to perform weights-based circuit analysis through MLP sublayers. The resulting circuits neatly factorize into input-dependent and input-invariant terms. We then successfully train transcoders on language models with 120M, 410M, and 1.4B parameters, and find them to perform at least on par with SAEs in terms of sparsity, faithfulness, and human-interpretability. Finally, we apply transcoders to reverse-engineer unknown circuits in the model, and we obtain novel insights regarding the "greater-than circuit" in GPT2-small. Our results suggest that transcoders can prove effective in decomposing model computations involving MLPs into interpretable circuits. Code is available at https://github.com/jacobdunefsky/transcoder_circuits/.

Language Models areFew-Shot LearnersLanguage Models are Few-Shot LearnersTowards AutomatedCircuit Discovery for…Towards Automated Circuit Discovery for Mechanistic InterpretabilityInterpretability in theWild: a Circuit for…Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallDoes Circuit AnalysisInterpretability Scale?…Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in ChinchillaFinding Neurons in aHaystack: Case Studies…Finding Neurons in a Haystack: Case Studies with Sparse ProbingLocalizing ModelBehavior with Path…Localizing Model Behavior with Path PatchingHow to use and interpretactivation patchingHow to use and interpret activation patchingSparse Autoencoders FindHighly Interpretable…Sparse Autoencoders Find Highly Interpretable Features in Language ModelsImproving DictionaryLearning with Gated…Improving Dictionary Learning with Gated Sparse AutoencodersTowards Best Practicesof Activation Patching…Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsDictionary LearningImproves Patch-Free…Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPTSparse Feature Circuits:Discovering and Editing…Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsLlama Scope: ExtractingMillions of Features…Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse AutoencodersInterPLM: DiscoveringInterpretable Features…InterPLM: Discovering Interpretable Features in Protein Language Models via Sparse AutoencodersMechanistic?Mechanistic?Sparse Autoencoders DoNot Find Canonical Unit…Sparse Autoencoders Do Not Find Canonical Units of AnalysisTranscoders Beat SparseAutoencoders for…Transcoders Beat Sparse Autoencoders for InterpretabilityResidual Stream Analysiswith Multi-Layer SAEsResidual Stream Analysis with Multi-Layer SAEsOn the TheoreticalFoundation of Sparse…On the Theoretical Foundation of Sparse Dictionary Learning in Mechanistic InterpretabilityWeight-sparsetransformers have…Weight-sparse transformers have interpretable circuitsPersona Vectors:Monitoring and…Persona Vectors: Monitoring and Controlling Character Traits in Language ModelsScaling sparse featurecircuit finding for…Scaling sparse feature circuit finding for in-context learningBridging the Black Box:A Survey on Mechanistic…Bridging the Black Box: A Survey on Mechanistic Interpretability in AIInterpretabilityIllusions with Sparse…Interpretability Illusions with Sparse Autoencoders: Evaluating Robustness of Concept RepresentationsTranscoders FindInterpretable LLM…Transcoders Find Interpretable LLM Feature Circuits過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。