Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Mechanistic interpretability seeks to understand the internal mechanisms of machine learning models, where localization -- identifying the important model components -- is a key step. Activation patching, also known as causal tracing or interchange intervention, is a standard technique for this task (Vig et al., 2020), but the literature contains many variants with little consensus on the choice of hyperparameters or methodology. In this work, we systematically examine the impact of methodological details in activation patching, including evaluation metrics and corruption methods. In several settings of localization and circuit discovery in language models, we find that varying these hyperparameters could lead to disparate interpretability results. Backed by empirical observations, we give conceptual arguments for why certain metrics or methods may be preferred. Finally, we provide recommendations for the best practices of activation patching going forwards.

Polysemanticity andCapacity in Neural…Polysemanticity and Capacity in Neural NetworksA MechanisticInterpretation of…A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation AnalysisLocalizing ModelBehavior with Path…Localizing Model Behavior with Path PatchingDoes Circuit AnalysisInterpretability Scale?…Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in ChinchillaThe Hydra Effect:Emergent Self-repair in…The Hydra Effect: Emergent Self-repair in Language Model ComputationsDissecting Recall ofFactual Associations in…Dissecting Recall of Factual Associations in Auto-Regressive Language ModelsEmergent LinearRepresentations in Worl…Emergent Linear Representations in World Models of Self-Supervised Sequence ModelsFinding Neurons in aHaystack: Case Studies…Finding Neurons in a Haystack: Case Studies with Sparse ProbingExplaining grokkingthrough circuit…Explaining grokking through circuit efficiencySparse Autoencoders FindHighly Interpretable…Sparse Autoencoders Find Highly Interpretable Features in Language ModelsFinding AlignmentsBetween Interpretable…Finding Alignments Between Interpretable Causal Variables and Distributed Neural RepresentationsLanguage ModelsImplement Simple…Language Models Implement Simple Word2Vec-style Vector ArithmeticHow to use and interpretactivation patchingHow to use and interpret activation patchingTranscoders FindInterpretable LLM…Transcoders Find Interpretable LLM Feature CircuitsDictionary LearningImproves Patch-Free…Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPTPatchscopes: A UnifyingFramework for Inspectin…Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsNeuron-Level KnowledgeAttribution in Large…Neuron-Level Knowledge Attribution in Large Language ModelsUnderstanding MultimodalLLMs: the Mechanistic…Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question AnsweringSparse AutoencodersEnable Scalable and…Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language ModelsRepresentationEngineering for…Representation Engineering for Large-Language Models: Survey and Research ChallengesNot All Language ModelFeatures Are…Not All Language Model Features Are One-Dimensionally LinearBack Attention:Understanding and…Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language ModelsHow Does Chain ofThought Think?…How Does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse AutoencodingLocate, Steer, andImprove: A Practical…Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language ModelsTowards Best Practicesof Activation Patching…Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.