Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Causal abstraction provides a theoretical foundation for mechanistic interpretability, the field concerned with providing intelligible algorithms that are faithful simplifications of the known, but opaque low-level details of black box AI models. Our contributions are (1) generalizing the theory of causal abstraction from mechanism replacement (i.e., hard and soft interventions) to arbitrary mechanism transformation (i.e., functionals from old mechanisms to new mechanisms), (2) providing a flexible, yet precise formalization for the core concepts of polysemantic neurons, the linear representation hypothesis, modular features, and graded faithfulness, and (3) unifying a variety of mechanistic interpretability methods in the common language of causal abstraction, namely, activation and path patching, causal mediation analysis, causal scrubbing, causal tracing, circuit analysis, concept erasure, sparse autoencoders, differential binary masking, distributed alignment search, and steering.

Neural Natural LanguageInference Models…Neural Natural Language Inference Models Partially Embed Theories of Lexical Entailment and NegationCausal Analysis ofSyntactic Agreement…Causal Analysis of Syntactic Agreement Mechanisms in Neural Language ModelsAn InterpretabilityIllusion for BERTAn Interpretability Illusion for BERTCEBaB: Estimating theCausal Effects of…CEBaB: Estimating the Causal Effects of Real-World Concepts on NLP Model BehaviorLocating and EditingFactual Associations in…Locating and Editing Factual Associations in GPTDoes Circuit AnalysisInterpretability Scale?…Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in ChinchillaCausal Proxy Models forConcept-Based Model…Causal Proxy Models for Concept-Based Model ExplanationsEmergent LinearRepresentations in Worl…Emergent Linear Representations in World Models of Self-Supervised Sequence ModelsFinding AlignmentsBetween Interpretable…Finding Alignments Between Interpretable Causal Variables and Distributed Neural RepresentationsReFT: RepresentationFinetuning for Language…ReFT: Representation Finetuning for Language ModelsMechanistic?Mechanistic?The LinearRepresentation…The Linear Representation Hypothesis and the Geometry of Large Language ModelsInterpretability atScale: Identifying…Interpretability at Scale: Identifying Causal Mechanisms in AlpacaDoes Circuit AnalysisInterpretability Scale?…Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in ChinchillaThe Hydra Effect:Emergent Self-repair in…The Hydra Effect: Emergent Self-repair in Language Model ComputationsIs This the Subspace YouAre Looking for? An…Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation PatchingA Primer on the InnerWorkings of…A Primer on the Inner Workings of Transformer-based Language ModelsReFT: RepresentationFinetuning for Language…ReFT: Representation Finetuning for Language ModelsMechanistic?Mechanistic?Combining Causal Modelsfor More Accurate…Combining Causal Models for More Accurate Abstractions of Neural NetworksMIB: A MechanisticInterpretability…MIB: A Mechanistic Interpretability BenchmarkAxBench: Steering LLMs?Even Simple Baselines…AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse AutoencodersA Survey on SparseAutoencoders…A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language ModelsSparse Feature Circuits:Discovering and Editing…Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsCausal Abstraction: ATheoretical Foundation…Causal Abstraction: A Theoretical Foundation for Mechanistic InterpretabilityEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.