Attribution Patching Outperforms Automated Circuit Discovery

Automated interpretability research has recently attracted attention as a potential research direction that could scale explanations of neural network behavior to large models. Existing automated circuit discovery work applies activation patching to identify subnetworks responsible for solving specific tasks (circuits). In this work, we show that a simple method based on attribution patching outperforms all existing methods while requiring just two forward passes and a backward pass. We apply a linear approximation to activation patching to estimate the importance of each edge in the computational subgraph. Using this approximation, we prune the least important edges of the network. We survey the performance and limitations of this method, finding that averaged over all tasks our method has greater AUC from circuit recovery than other methods.

Causal MediationAnalysis for…Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender BiasAn overview of 11proposals for building…An overview of 11 proposals for building safe advanced AICausal Abstractions ofNeural NetworksCausal Abstractions of Neural NetworksLow-Complexity Probingvia Finding SubnetworksLow-Complexity Probing via Finding SubnetworksHow does GPT-2 computegreater-than?…How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language modelDoes Circuit AnalysisInterpretability Scale?…Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in ChinchillaA Survey on ModelCompression for Large…A Survey on Model Compression for Large Language ModelsTowards AutomatedCircuit Discovery for…Towards Automated Circuit Discovery for Mechanistic InterpretabilityHave Faith inFaithfulness: Going…Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model MechanismsSparse AutoencodersEnable Scalable and…Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language ModelsA Practical Review ofMechanistic…A Practical Review of Mechanistic Interpretability for Transformer-Based Language ModelsA Primer on the InnerWorkings of…A Primer on the Inner Workings of Transformer-based Language ModelsInversionView: AGeneral-Purpose Method…InversionView: A General-Purpose Method for Reading Information from Neural ActivationsDecomposing and EditingPredictions by Modeling…Decomposing and Editing Predictions by Modeling Model ComputationUnboxing the Black Box:Mechanistic…Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural NetworksEfficient AutomatedCircuit Discovery in…Efficient Automated Circuit Discovery in Transformers using Contextual DecompositionEAP-GP: MitigatingSaturation Effect in…EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit IdentificationInterpreting visiontransformers via…Interpreting vision transformers via residual replacement modelDiscovering TransformerCircuits via a Hybrid…Discovering Transformer Circuits via a Hybrid Attribution and Pruning FrameworkAttribution PatchingOutperforms Automated…Attribution Patching Outperforms Automated Circuit Discovery過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。