著者: Javier Ferrando , Gabriele Sarti , Arianna Bisazza , Marta R. Costa-jussà - arXiv 2024 被引用: 120
The rapid progress of research aimed at interpreting the inner workings of advanced language models has highlighted a need for contextualizing the insights gained from years of work in this area. This primer provides a concise technical introduction to the current techniques used to interpret the inner workings of Transformer-based language models, focusing on the generative decoder-only architecture. We conclude by presenting a comprehensive overview of the known internal mechanisms implemented by these models, uncovering connections across popular approaches and active research directions in this area.
✨ ログイン状態を確認しています… PDF 被引用 BibTeX を表示 BibTeX を閉じる BibTeX を表示 引用
Incorporating Residual and Normalization Layer… Incorporating Residual and Normalization Layers into Analysis of Masked Language Models Post-hoc Interpretability for… Post-hoc Interpretability for Neural NLP: A Survey Dissecting Recall of Factual Associations in… Dissecting Recall of Factual Associations in Auto-Regressive Language Models Does Circuit Analysis Interpretability Scale?… Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla Refusal in Language Models Is Mediated by a… Refusal in Language Models Is Mediated by a Single Direction Information Flow Routes: Automatically… Information Flow Routes: Automatically Interpreting Language Models at Scale How to use and interpret activation patching How to use and interpret activation patching The Linear Representation… The Linear Representation Hypothesis and the Geometry of Large Language Models AtP*: An efficient and scalable method for… AtP*: An efficient and scalable method for localizing LLM behaviour to components Decomposing and Editing Predictions by Modeling… Decomposing and Editing Predictions by Modeling Model Computation Towards Principled Evaluations of Sparse… Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control Sparse Feature Circuits: Discovering and Editing… Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models A Practical Review of Mechanistic… A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models Mechanistic? Mechanistic? Transcoders Find Interpretable LLM… Transcoders Find Interpretable LLM Feature Circuits The Remarkable Robustness of LLMs… The Remarkable Robustness of LLMs: Stages of Inference? Usable XAI: 10 Strategies Towards… Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era Measuring Progress in Dictionary Learning for… Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models A Survey on Sparse Autoencoders… A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models Representation Engineering for… Representation Engineering for Large-Language Models: Survey and Research Challenges Interpreting vision transformers via… Interpreting vision transformers via residual replacement model Interpreting Attention Heads for Image-to-Text… Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models Steering off Course: Reliability Challenges… Steering off Course: Reliability Challenges in Steering Language Models Locate, Steer, and Improve: A Practical… Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models A Primer on the Inner Workings of… A Primer on the Inner Workings of Transformer-based Language Models 過去の参考文献 中心の論文 この論文を引用する論文 古い 新しい ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。