Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small

Research in mechanistic interpretability seeks to explain behaviors of machine learning models in terms of their internal components. However, most previous work either focuses on simple behaviors in small models, or describes complicated behaviors in larger models with broad strokes. In this work, we bridge this gap by presenting an explanation for how GPT-2 small performs a natural language task called indirect object identification (IOI). Our explanation encompasses 26 attention heads grouped into 7 main classes, which we discovered using a combination of interpretability approaches relying on causal interventions. To our knowledge, this investigation is the largest end-to-end attempt at reverse-engineering a natural behavior "in the wild" in a language model. We evaluate the reliability of our explanation using three quantitative criteria--faithfulness, completeness and minimality. Though these criteria support our explanation, they also point to remaining gaps in our understanding. Our work provides evidence that a mechanistic understanding of large ML models is feasible, opening opportunities to scale our understanding to both larger models and more complex tasks.

Attention is notExplanationAttention is not ExplanationTransformer Feed-ForwardLayers Are Key-Value…Transformer Feed-Forward Layers Are Key-Value MemoriesAn InterpretabilityIllusion for BERTAn Interpretability Illusion for BERTLocating and EditingFactual Associations in…Locating and Editing Factual Associations in GPTIn-context Learning andInduction HeadsIn-context Learning and Induction HeadsHidden Progress in DeepLearning: SGD Learns…Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitToward Transparent AI: ASurvey on Interpreting…Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural NetworksAttentionViz: A GlobalView of Transformer…AttentionViz: A Global View of Transformer AttentionDoes Circuit AnalysisInterpretability Scale?…Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in ChinchillaA MechanisticInterpretation of…A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation AnalysisEliciting LatentPredictions from…Eliciting Latent Predictions from Transformers with the Tuned LensTransformers LearnShortcuts to AutomataTransformers Learn Shortcuts to AutomataSparse Autoencoders FindHighly Interpretable…Sparse Autoencoders Find Highly Interpretable Features in Language ModelsChain of ThoughtEmpowers Transformers t…Chain of Thought Empowers Transformers to Solve Inherently Serial ProblemsCodebook Features:Sparse and Discrete…Codebook Features: Sparse and Discrete Interpretability for Neural NetworksBack Attention:Understanding and…Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language ModelsSoftmax is not Enough(for Sharp Size…Softmax is not Enough (for Sharp Size Generalisation)Physics of LanguageModels: Part 1, Learnin…Physics of Language Models: Part 1, Learning Hierarchical Language StructuresWhen Truth IsOverridden: Uncovering…When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language ModelsInterpretability in theWild: a Circuit for…Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.