著者: Stefan Heimersheim , Neel Nanda - arXiv 2024 被引用: 118
Activation patching is a popular mechanistic interpretability technique, but has many subtleties regarding how it is applied and how one may interpret the results. We provide a summary of advice and best practices, based on our experience using this technique in practice. We include an overview of the different ways to apply activation patching and a discussion on how to interpret the results. We focus on what evidence patching experiments provide about circuits, and on the choice of metric and associated pitfalls.
✨ ログイン状態を確認しています… PDF 被引用 BibTeX を表示 BibTeX を閉じる BibTeX を表示 引用
Causal Abstractions of Neural Networks Causal Abstractions of Neural Networks Causal Analysis of Syntactic Agreement… Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models Locating and Editing Factual Associations in… Locating and Editing Factual Associations in GPT Towards Automated Circuit Discovery for… Towards Automated Circuit Discovery for Mechanistic Interpretability Does Circuit Analysis Interpretability Scale?… Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla The Hydra Effect: Emergent Self-repair in… The Hydra Effect: Emergent Self-repair in Language Model Computations Interpretability in the Wild: a Circuit for… Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small Localizing Model Behavior with Path… Localizing Model Behavior with Path Patching Dissecting Recall of Factual Associations in… Dissecting Recall of Factual Associations in Auto-Regressive Language Models Copy Suppression: Comprehensively… Copy Suppression: Comprehensively Understanding an Attention Head A Mechanistic Interpretation of… A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis Towards Best Practices of Activation Patching… Towards Best Practices of Activation Patching in Language Models: Metrics and Methods Transcoders Find Interpretable LLM… Transcoders Find Interpretable LLM Feature Circuits A Practical Review of Mechanistic… A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models A Primer on the Inner Workings of… A Primer on the Inner Workings of Transformer-based Language Models Missed Causes and Ambiguous Effects… Missed Causes and Ambiguous Effects: Counterfactuals Pose Challenges for Interpreting Neural Networks Attention Heads of Large Language Models: A… Attention Heads of Large Language Models: A Survey Addressing divergent representations from… Addressing divergent representations from causal interventions on neural networks Unboxing the Black Box: Mechanistic… Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks How do Large Language Models Understand… How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective Do I Know This Entity? Knowledge Awareness and… Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models Aligned but Blind: Alignment Increases… Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Representation Engineering for… Representation Engineering for Large-Language Models: Survey and Research Challenges Rethinking Circuit Completeness in Languag… Rethinking Circuit Completeness in Language Models: AND, OR, and ADDER Gates How to use and interpret activation patching How to use and interpret activation patching 過去の参考文献 中心の論文 この論文を引用する論文 古い 新しい ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。