The Hydra Effect: Emergent Self-repair in Language Model Computations

We investigate the internal structure of language model computations using causal analysis and demonstrate two motifs: (1) a form of adaptive computation where ablations of one attention layer of a language model cause another layer to compensate (which we term the Hydra effect) and (2) a counterbalancing function of late MLP layers that act to downregulate the maximum-likelihood token. Our ablation studies demonstrate that language model layers are typically relatively loosely coupled (ablations to one layer only affect a small number of downstream layers). Surprisingly, these effects occur even in language models trained without any form of dropout. We analyse these effects in the context of factual recall and consider their implications for circuit-level attribution in language models.

Neural MachineTranslation by Jointly…Neural Machine Translation by Jointly Learning to Align and TranslateUnderstandingintermediate layers…Understanding intermediate layers using linear classifier probesCausal Analysis ofSyntactic Agreement…Causal Analysis of Syntactic Agreement Mechanisms in Neural Language ModelsIn-context Learning andInduction HeadsIn-context Learning and Induction HeadsTowards AutomatedCircuit Discovery for…Towards Automated Circuit Discovery for Mechanistic InterpretabilityFinding Neurons in aHaystack: Case Studies…Finding Neurons in a Haystack: Case Studies with Sparse ProbingInterpretability in theWild: a Circuit for…Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallEliciting LatentPredictions from…Eliciting Latent Predictions from Transformers with the Tuned LensProgress measures forgrokking via mechanisti…Progress measures for grokking via mechanistic interpretabilityAnalyzing Transformersin Embedding SpaceAnalyzing Transformers in Embedding SpaceFinding AlignmentsBetween Interpretable…Finding Alignments Between Interpretable Causal Variables and Distributed Neural RepresentationsCausal Abstraction: ATheoretical Foundation…Causal Abstraction: A Theoretical Foundation for Mechanistic InterpretabilityEmergent LinearRepresentations in Worl…Emergent Linear Representations in World Models of Self-Supervised Sequence ModelsLearning TransformerProgramsLearning Transformer ProgramsRepresentationEngineering: A Top-Down…Representation Engineering: A Top-Down Approach to AI TransparencyHow to use and interpretactivation patchingHow to use and interpret activation patchingUniversal Neurons inGPT2 Language ModelsUniversal Neurons in GPT2 Language ModelsTowards Best Practicesof Activation Patching…Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsSumming Up the Facts:Additive Mechanisms…Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMsThe RemarkableRobustness of LLMs…The Remarkable Robustness of LLMs: Stages of Inference?Is This the Subspace YouAre Looking for? An…Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation PatchingAutomaticallyIdentifying Local and…Automatically Identifying Local and Global Circuits with Linear Computation GraphsTowards PrincipledEvaluations of Sparse…Towards Principled Evaluations of Sparse Autoencoders for Interpretability and ControlOpen Problems inMechanistic…Open Problems in Mechanistic InterpretabilityThe Hydra Effect:Emergent Self-repair in…The Hydra Effect: Emergent Self-repair in Language Model Computations過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。