著者: Thomas McGrath , Matthew Rahtz , János Kramár , Vladimir Mikulik , Shane Legg - arXiv 2023 被引用: 75
We investigate the internal structure of language model computations using causal analysis and demonstrate two motifs: (1) a form of adaptive computation where ablations of one attention layer of a language model cause another layer to compensate (which we term the Hydra effect) and (2) a counterbalancing function of late MLP layers that act to downregulate the maximum-likelihood token. Our ablation studies demonstrate that language model layers are typically relatively loosely coupled (ablations to one layer only affect a small number of downstream layers). Surprisingly, these effects occur even in language models trained without any form of dropout. We analyse these effects in the context of factual recall and consider their implications for circuit-level attribution in language models.
✨ ログイン状態を確認しています… PDF 被引用 BibTeX を表示 BibTeX を閉じる BibTeX を表示 引用
Neural Machine Translation by Jointly… Neural Machine Translation by Jointly Learning to Align and Translate Understanding intermediate layers… Understanding intermediate layers using linear classifier probes Causal Analysis of Syntactic Agreement… Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models In-context Learning and Induction Heads In-context Learning and Induction Heads Towards Automated Circuit Discovery for… Towards Automated Circuit Discovery for Mechanistic Interpretability Finding Neurons in a Haystack: Case Studies… Finding Neurons in a Haystack: Case Studies with Sparse Probing Interpretability in the Wild: a Circuit for… Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small Eliciting Latent Predictions from… Eliciting Latent Predictions from Transformers with the Tuned Lens Progress measures for grokking via mechanisti… Progress measures for grokking via mechanistic interpretability Analyzing Transformers in Embedding Space Analyzing Transformers in Embedding Space Finding Alignments Between Interpretable… Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations Causal Abstraction: A Theoretical Foundation… Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability Emergent Linear Representations in Worl… Emergent Linear Representations in World Models of Self-Supervised Sequence Models Learning Transformer Programs Learning Transformer Programs Representation Engineering: A Top-Down… Representation Engineering: A Top-Down Approach to AI Transparency How to use and interpret activation patching How to use and interpret activation patching Universal Neurons in GPT2 Language Models Universal Neurons in GPT2 Language Models Towards Best Practices of Activation Patching… Towards Best Practices of Activation Patching in Language Models: Metrics and Methods Summing Up the Facts: Additive Mechanisms… Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs The Remarkable Robustness of LLMs… The Remarkable Robustness of LLMs: Stages of Inference? Is This the Subspace You Are Looking for? An… Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching Automatically Identifying Local and… Automatically Identifying Local and Global Circuits with Linear Computation Graphs Towards Principled Evaluations of Sparse… Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control Open Problems in Mechanistic… Open Problems in Mechanistic Interpretability The Hydra Effect: Emergent Self-repair in… The Hydra Effect: Emergent Self-repair in Language Model Computations 過去の参考文献 中心の論文 この論文を引用する論文 古い 新しい ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。