AtP*: An efficient and scalable method for localizing LLM behaviour to components

Activation Patching is a method of directly computing causal attributions of behavior to model components. However, applying it exhaustively requires a sweep with cost scaling linearly in the number of model components, which can be prohibitively expensive for SoTA Large Language Models (LLMs). We investigate Attribution Patching (AtP), a fast gradient-based approximation to Activation Patching and find two classes of failure modes of AtP which lead to significant false negatives. We propose a variant of AtP called AtP*, with two changes to address these failure modes while retaining scalability. We present the first systematic study of AtP and alternative methods for faster activation patching and show that AtP significantly outperforms all other investigated methods, with AtP* providing further significant improvement. Finally, we provide a method to bound the probability of remaining false negatives of AtP* estimates.

The Pile: An 800GBDataset of Diverse Text…The Pile: An 800GB Dataset of Diverse Text for Language ModelingDiscovering theCompositional Structure…Discovering the Compositional Structure of Vector Representations with Role Learning NetworksCausal Analysis ofSyntactic Agreement…Causal Analysis of Syntactic Agreement Mechanisms in Neural Language ModelsLEACE: Perfect linearconcept erasure in…LEACE: Perfect linear concept erasure in closed formThe Quest for the RightMediator: A History…The Quest for the Right Mediator: A History, Survey, and Theoretical Grounding of Causal InterpretabilityMechanistic?Mechanistic?A Primer on the InnerWorkings of…A Primer on the Inner Workings of Transformer-based Language ModelsMissed Causes andAmbiguous Effects…Missed Causes and Ambiguous Effects: Counterfactuals Pose Challenges for Interpreting Neural NetworksDecomposing and EditingPredictions by Modeling…Decomposing and Editing Predictions by Modeling Model ComputationSparse Feature Circuits:Discovering and Editing…Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsStochastic ParameterDecompositionStochastic Parameter DecompositionOpen Problems inMechanistic…Open Problems in Mechanistic InterpretabilityNNsight and NDIF:Democratizing Access to…NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model InternalsInterpretability inParameter Space…Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter DecompositionPosition-aware AutomaticCircuit DiscoveryPosition-aware Automatic Circuit DiscoveryInterpreting visiontransformers via…Interpreting vision transformers via residual replacement modelAtP*: An efficient andscalable method for…AtP*: An efficient and scalable method for localizing LLM behaviour to components過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。