Authors: Javier Ferrando , Gabriele Sarti , Arianna Bisazza , Marta R. Costa-jussà - arXiv 2024 cited by 120
The rapid progress of research aimed at interpreting the inner workings of advanced language models has highlighted a need for contextualizing the insights gained from years of work in this area. This primer provides a concise technical introduction to the current techniques used to interpret the inner workings of Transformer-based language models, focusing on the generative decoder-only architecture. We conclude by presenting a comprehensive overview of the known internal mechanisms implemented by these models, uncovering connections across popular approaches and active research directions in this area.
✨ Checking sign-in… PDF Cited by View BibTeX Hide BibTeX View BibTeX Cite
Incorporating Residual and Normalization Layer… Incorporating Residual and Normalization Layers into Analysis of Masked Language Models Post-hoc Interpretability for… Post-hoc Interpretability for Neural NLP: A Survey Dissecting Recall of Factual Associations in… Dissecting Recall of Factual Associations in Auto-Regressive Language Models Does Circuit Analysis Interpretability Scale?… Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla Refusal in Language Models Is Mediated by a… Refusal in Language Models Is Mediated by a Single Direction Information Flow Routes: Automatically… Information Flow Routes: Automatically Interpreting Language Models at Scale How to use and interpret activation patching How to use and interpret activation patching The Linear Representation… The Linear Representation Hypothesis and the Geometry of Large Language Models AtP*: An efficient and scalable method for… AtP*: An efficient and scalable method for localizing LLM behaviour to components Decomposing and Editing Predictions by Modeling… Decomposing and Editing Predictions by Modeling Model Computation Towards Principled Evaluations of Sparse… Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control Sparse Feature Circuits: Discovering and Editing… Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models A Practical Review of Mechanistic… A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models Mechanistic? Mechanistic? Transcoders Find Interpretable LLM… Transcoders Find Interpretable LLM Feature Circuits The Remarkable Robustness of LLMs… The Remarkable Robustness of LLMs: Stages of Inference? Usable XAI: 10 Strategies Towards… Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era Measuring Progress in Dictionary Learning for… Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models A Survey on Sparse Autoencoders… A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models Representation Engineering for… Representation Engineering for Large-Language Models: Survey and Research Challenges Interpreting vision transformers via… Interpreting vision transformers via residual replacement model Interpreting Attention Heads for Image-to-Text… Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models Steering off Course: Reliability Challenges… Steering off Course: Reliability Challenges in Steering Language Models Locate, Steer, and Improve: A Practical… Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models A Primer on the Inner Workings of… A Primer on the Inner Workings of Transformer-based Language Models Earlier references Focus paper Citing papers Older Newer Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.