Authors: Thomas McGrath , Matthew Rahtz , János Kramár , Vladimir Mikulik , Shane Legg - arXiv 2023 cited by 75
We investigate the internal structure of language model computations using causal analysis and demonstrate two motifs: (1) a form of adaptive computation where ablations of one attention layer of a language model cause another layer to compensate (which we term the Hydra effect) and (2) a counterbalancing function of late MLP layers that act to downregulate the maximum-likelihood token. Our ablation studies demonstrate that language model layers are typically relatively loosely coupled (ablations to one layer only affect a small number of downstream layers). Surprisingly, these effects occur even in language models trained without any form of dropout. We analyse these effects in the context of factual recall and consider their implications for circuit-level attribution in language models.
✨ Checking sign-in… PDF Cited by View BibTeX Hide BibTeX View BibTeX Cite
Neural Machine Translation by Jointly… Neural Machine Translation by Jointly Learning to Align and Translate Understanding intermediate layers… Understanding intermediate layers using linear classifier probes Causal Analysis of Syntactic Agreement… Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models In-context Learning and Induction Heads In-context Learning and Induction Heads Towards Automated Circuit Discovery for… Towards Automated Circuit Discovery for Mechanistic Interpretability Finding Neurons in a Haystack: Case Studies… Finding Neurons in a Haystack: Case Studies with Sparse Probing Interpretability in the Wild: a Circuit for… Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small Eliciting Latent Predictions from… Eliciting Latent Predictions from Transformers with the Tuned Lens Progress measures for grokking via mechanisti… Progress measures for grokking via mechanistic interpretability Analyzing Transformers in Embedding Space Analyzing Transformers in Embedding Space Finding Alignments Between Interpretable… Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations Causal Abstraction: A Theoretical Foundation… Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability Emergent Linear Representations in Worl… Emergent Linear Representations in World Models of Self-Supervised Sequence Models Learning Transformer Programs Learning Transformer Programs Representation Engineering: A Top-Down… Representation Engineering: A Top-Down Approach to AI Transparency How to use and interpret activation patching How to use and interpret activation patching Universal Neurons in GPT2 Language Models Universal Neurons in GPT2 Language Models Towards Best Practices of Activation Patching… Towards Best Practices of Activation Patching in Language Models: Metrics and Methods Summing Up the Facts: Additive Mechanisms… Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs The Remarkable Robustness of LLMs… The Remarkable Robustness of LLMs: Stages of Inference? Is This the Subspace You Are Looking for? An… Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching Automatically Identifying Local and… Automatically Identifying Local and Global Circuits with Linear Computation Graphs Towards Principled Evaluations of Sparse… Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control Open Problems in Mechanistic… Open Problems in Mechanistic Interpretability The Hydra Effect: Emergent Self-repair in… The Hydra Effect: Emergent Self-repair in Language Model Computations Earlier references Focus paper Citing papers Older Newer Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.