Localizing Model Behavior with Path Patching

Localizing behaviors of neural networks to a subset of the network's components or a subset of interactions between components is a natural first step towards analyzing network mechanisms and possible failure modes. Existing work is often qualitative and ad-hoc, and there is no consensus on the appropriate way to evaluate localization claims. We introduce path patching, a technique for expressing and quantitatively testing a natural class of hypotheses expressing that behaviors are localized to a set of paths. We refine an explanation of induction heads, characterize a behavior of GPT-2, and open source a framework for efficiently running similar experiments.

Modular RepresentationUnderlies Systematic…Modular Representation Underlies Systematic Generalization in Neural Natural Language Inference ModelsAn InterpretabilityIllusion for BERTAn Interpretability Illusion for BERTCausal Analysis ofSyntactic Agreement…Causal Analysis of Syntactic Agreement Mechanisms in Neural Language ModelsShortformer: BetterLanguage Modeling using…Shortformer: Better Language Modeling using Shorter InputsFinding AlignmentsBetween Interpretable…Finding Alignments Between Interpretable Causal Variables and Distributed Neural RepresentationsCausal Abstraction: ATheoretical Foundation…Causal Abstraction: A Theoretical Foundation for Mechanistic InterpretabilityTowards AutomatedCircuit Discovery for…Towards Automated Circuit Discovery for Mechanistic InterpretabilityFinding Neurons in aHaystack: Case Studies…Finding Neurons in a Haystack: Case Studies with Sparse ProbingTowards Best Practicesof Activation Patching…Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsTranscoders FindInterpretable LLM…Transcoders Find Interpretable LLM Feature CircuitsHow to use and interpretactivation patchingHow to use and interpret activation patchingCircuit Component ReuseAcross Tasks in…Circuit Component Reuse Across Tasks in Transformer Language ModelsIs This the Subspace YouAre Looking for? An…Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation PatchingPatchscopes: A UnifyingFramework for Inspectin…Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsDecomposing and EditingPredictions by Modeling…Decomposing and Editing Predictions by Modeling Model ComputationTalking Heads:Understanding…Talking Heads: Understanding Inter-Layer Communication in Transformer Language ModelsNeuron-Level KnowledgeAttribution in Large…Neuron-Level Knowledge Attribution in Large Language ModelsAxBench: Steering LLMs?Even Simple Baselines…AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse AutoencodersLocalizing ModelBehavior with Path…Localizing Model Behavior with Path Patching過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。