Open Problems in Mechanistic Interpretability

Mechanistic interpretability aims to understand the computational mechanisms underlying neural networks' capabilities in order to accomplish concrete scientific and engineering goals. Progress in this field thus promises to provide greater assurance over AI system behavior and shed light on exciting scientific questions about the nature of intelligence. Despite recent progress toward these goals, there are many open problems in the field that require solutions before many scientific and practical benefits can be realized: Our methods require both conceptual and practical improvements to reveal deeper insights; we must figure out how best to apply our methods in pursuit of specific goals; and the field must grapple with socio-technical challenges that influence and are influenced by our work. This forward-facing review discusses the current frontier of mechanistic interpretability and the open problems that the field may benefit from prioritizing.

Deep Residual Learningfor Image RecognitionDeep Residual Learning for Image RecognitionAn InterpretabilityIllusion for BERTAn Interpretability Illusion for BERTTransformervisualization via…Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factorsLocating and EditingFactual Associations in…Locating and Editing Factual Associations in GPTDoes Circuit AnalysisInterpretability Scale?…Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in ChinchillaDissecting Recall ofFactual Associations in…Dissecting Recall of Factual Associations in Auto-Regressive Language ModelsRepresentationEngineering: A Top-Down…Representation Engineering: A Top-Down Approach to AI TransparencyMechanistic?Mechanistic?Is This the Subspace YouAre Looking for? An…Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation PatchingSparse Feature Circuits:Discovering and Editing…Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsA is for Absorption:Studying Feature…A is for Absorption: Studying Feature Splitting and Absorption in Sparse AutoencodersNot All Language ModelFeatures Are…Not All Language Model Features Are One-Dimensionally LinearAre Sparse AutoencodersUseful? A Case Study in…Are Sparse Autoencoders Useful? A Case Study in Sparse ProbingThe Geometry ofConcepts: Sparse…The Geometry of Concepts: Sparse Autoencoder Feature StructureInterpretation MeetsSafety: A Survey on…Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM SafetyInterpreting visiontransformers via…Interpreting vision transformers via residual replacement modelChain of ThoughtMonitorability: A New…Chain of Thought Monitorability: A New and Fragile Opportunity for AI SafetyActivation SpaceInterventions Can Be…Activation Space Interventions Can Be Transferred Between Large Language ModelsYou Are What You Eat -AI Alignment Requires…You Are What You Eat - AI Alignment Requires Understanding How Data Shapes Structure and GeneralisationAgentic Large LanguageModels, a SurveyAgentic Large Language Models, a SurveyGroup-SAE: EfficientTraining of Sparse…Group-SAE: Efficient Training of Sparse Autoencoders for Large Language Models via Layer GroupsLow-Rank Adapting Modelsfor Sparse AutoencodersLow-Rank Adapting Models for Sparse AutoencodersThe Information Geometryof Softmax: Probing and…The Information Geometry of Softmax: Probing and SteeringInterpretabilityIllusions with Sparse…Interpretability Illusions with Sparse Autoencoders: Evaluating Robustness of Concept RepresentationsOpen Problems inMechanistic…Open Problems in Mechanistic InterpretabilityEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.