A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models

Mechanistic interpretability (MI) is an emerging sub-field of interpretability that seeks to understand a neural network model by reverse-engineering its internal computations. Recently, MI has garnered significant attention for interpreting transformer-based language models (LMs), resulting in many novel insights yet introducing new challenges. However, there has not been work that comprehensively reviews these insights and challenges, particularly as a guide for newcomers to this field. To fill this gap, we provide a comprehensive survey from a task-centric perspective, organizing the taxonomy of MI research around specific research questions or tasks. We outline the fundamental objects of study in MI, along with the techniques, evaluation methods, and key findings for each task in the taxonomy. In particular, we present a task-centric taxonomy as a roadmap for beginners to navigate the field by helping them quickly identify impactful problems in which they are most interested and leverage MI for their benefit. Finally, we discuss the current gaps in the field and suggest potential future directions for MI research.

Does Circuit AnalysisInterpretability Scale?…Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in ChinchillaA MechanisticInterpretation of…A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation AnalysisMechanistic?Mechanistic?Universal Neurons inGPT2 Language ModelsUniversal Neurons in GPT2 Language ModelsA Primer on the InnerWorkings of…A Primer on the Inner Workings of Transformer-based Language ModelsLlama Scope: ExtractingMillions of Features…Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse AutoencodersRefusal in LanguageModels Is Mediated by a…Refusal in Language Models Is Mediated by a Single DirectionHow to use and interpretactivation patchingHow to use and interpret activation patchingA Survey on SparseAutoencoders…A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language ModelsNot All Language ModelFeatures Are…Not All Language Model Features Are One-Dimensionally LinearOpen Problems inMechanistic…Open Problems in Mechanistic InterpretabilitySparse Feature Circuits:Discovering and Editing…Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsUniversal Response andEmergence of Induction…Universal Response and Emergence of Induction in LLMsA Survey on SparseAutoencoders…A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language ModelsInterpreting visiontransformers via…Interpreting vision transformers via residual replacement modelRepresentationEngineering for…Representation Engineering for Large-Language Models: Survey and Research ChallengesBuilding Bridges, NotWalls - Advancing…Building Bridges, Not Walls - Advancing Interpretability by Unifying Feature, Data, and Model Component AttributionSteering off Course:Reliability Challenges…Steering off Course: Reliability Challenges in Steering Language ModelsBridging the Black Box:A Survey on Mechanistic…Bridging the Black Box: A Survey on Mechanistic Interpretability in AIUnpacking SDXL Turbo:Interpreting…Unpacking SDXL Turbo: Interpreting Text-to-Image Models with Sparse AutoencodersUnderstanding the RepeatCurse in Large Language…Understanding the Repeat Curse in Large Language Models from a Feature PerspectiveTowards UnderstandingFine-Tuning Mechanisms…Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit AnalysisAdaptiveK:Complexity-Driven Spars…AdaptiveK: Complexity-Driven Sparse Autoencoders for Interpretable Language Model RepresentationsLocate, Steer, andImprove: A Practical…Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language ModelsA Practical Review ofMechanistic…A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。