Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla

\emph{Circuit analysis} is a promising technique for understanding the internal mechanisms of language models. However, existing analyses are done in small models far from the state of the art. To address this, we present a case study of circuit analysis in the 70B Chinchilla model, aiming to test the scalability of circuit analysis. In particular, we study multiple-choice question answering, and investigate Chinchilla's capability to identify the correct answer \emph{label} given knowledge of the correct answer \emph{text}. We find that the existing techniques of logit attribution, attention pattern visualization, and activation patching naturally scale to Chinchilla, allowing us to identify and categorize a small set of `output nodes' (attention heads and MLPs). We further study the `correct letter' category of attention heads aiming to understand the semantics of their features, with mixed results. For normal multiple-choice question answers, we significantly compress the query, key and value subspaces of the head without loss of performance when operating on the answer labels for multiple-choice questions, and we show that the query and key subspaces represent an `Nth item in an enumeration' feature to at least some extent. However, when we attempt to use this explanation to understand the heads' behaviour on a more general distribution including randomized answer labels, we find that it is only a partial explanation, suggesting there is more to learn about the operation of `correct letter' heads on multiple choice question answering.

Improving alignment ofdialogue agents via…Improving alignment of dialogue agents via targeted human judgementsTowards AutomatedCircuit Discovery for…Towards Automated Circuit Discovery for Mechanistic InterpretabilityInterpretability in theWild: a Circuit for…Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallDissecting Recall ofFactual Associations in…Dissecting Recall of Factual Associations in Auto-Regressive Language ModelsFinding Neurons in aHaystack: Case Studies…Finding Neurons in a Haystack: Case Studies with Sparse ProbingInterpretability atScale: Identifying…Interpretability at Scale: Identifying Causal Mechanisms in AlpacaEliciting LatentPredictions from…Eliciting Latent Predictions from Transformers with the Tuned LensProgress measures forgrokking via mechanisti…Progress measures for grokking via mechanistic interpretabilityAnalyzing Transformersin Embedding SpaceAnalyzing Transformers in Embedding SpaceFinding AlignmentsBetween Interpretable…Finding Alignments Between Interpretable Causal Variables and Distributed Neural RepresentationsLet's Verify Step byStepLet's Verify Step by StepCausal Abstraction: ATheoretical Foundation…Causal Abstraction: A Theoretical Foundation for Mechanistic InterpretabilityHow to use and interpretactivation patchingHow to use and interpret activation patchingAttribution PatchingOutperforms Automated…Attribution Patching Outperforms Automated Circuit DiscoveryTranscoders FindInterpretable LLM…Transcoders Find Interpretable LLM Feature CircuitsA Practical Review ofMechanistic…A Practical Review of Mechanistic Interpretability for Transformer-Based Language ModelsTowards Best Practicesof Activation Patching…Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsHave Faith inFaithfulness: Going…Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model MechanismsNeuron-Level KnowledgeAttribution in Large…Neuron-Level Knowledge Attribution in Large Language ModelsUniversal Response andEmergence of Induction…Universal Response and Emergence of Induction in LLMsDictionary LearningImproves Patch-Free…Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPTSparse AutoencodersEnable Scalable and…Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language ModelsInterpreting visiontransformers via…Interpreting vision transformers via residual replacement modelAligned but Blind:Alignment Increases…Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of RaceDoes Circuit AnalysisInterpretability Scale?…Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in ChinchillaEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.