A Primer on the Inner Workings of Transformer-based Language Models

The rapid progress of research aimed at interpreting the inner workings of advanced language models has highlighted a need for contextualizing the insights gained from years of work in this area. This primer provides a concise technical introduction to the current techniques used to interpret the inner workings of Transformer-based language models, focusing on the generative decoder-only architecture. We conclude by presenting a comprehensive overview of the known internal mechanisms implemented by these models, uncovering connections across popular approaches and active research directions in this area.

Incorporating Residualand Normalization Layer…Incorporating Residual and Normalization Layers into Analysis of Masked Language ModelsPost-hocInterpretability for…Post-hoc Interpretability for Neural NLP: A SurveyDissecting Recall ofFactual Associations in…Dissecting Recall of Factual Associations in Auto-Regressive Language ModelsDoes Circuit AnalysisInterpretability Scale?…Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in ChinchillaRefusal in LanguageModels Is Mediated by a…Refusal in Language Models Is Mediated by a Single DirectionInformation Flow Routes:Automatically…Information Flow Routes: Automatically Interpreting Language Models at ScaleHow to use and interpretactivation patchingHow to use and interpret activation patchingThe LinearRepresentation…The Linear Representation Hypothesis and the Geometry of Large Language ModelsAtP*: An efficient andscalable method for…AtP*: An efficient and scalable method for localizing LLM behaviour to componentsDecomposing and EditingPredictions by Modeling…Decomposing and Editing Predictions by Modeling Model ComputationTowards PrincipledEvaluations of Sparse…Towards Principled Evaluations of Sparse Autoencoders for Interpretability and ControlSparse Feature Circuits:Discovering and Editing…Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsA Practical Review ofMechanistic…A Practical Review of Mechanistic Interpretability for Transformer-Based Language ModelsMechanistic?Mechanistic?Transcoders FindInterpretable LLM…Transcoders Find Interpretable LLM Feature CircuitsThe RemarkableRobustness of LLMs…The Remarkable Robustness of LLMs: Stages of Inference?Usable XAI: 10Strategies Towards…Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM EraMeasuring Progress inDictionary Learning for…Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game ModelsA Survey on SparseAutoencoders…A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language ModelsRepresentationEngineering for…Representation Engineering for Large-Language Models: Survey and Research ChallengesInterpreting visiontransformers via…Interpreting vision transformers via residual replacement modelInterpreting AttentionHeads for Image-to-Text…Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language ModelsSteering off Course:Reliability Challenges…Steering off Course: Reliability Challenges in Steering Language ModelsLocate, Steer, andImprove: A Practical…Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language ModelsA Primer on the InnerWorkings of…A Primer on the Inner Workings of Transformer-based Language ModelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.