著者: Aaquib Syed , Can Rager , Arthur Conmy - Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP 2024 被引用: 70
Automated interpretability research has recently attracted attention as a potential research direction that could scale explanations of neural network behavior to large models. Existing automated circuit discovery work applies activation patching to identify subnetworks responsible for solving specific tasks (circuits). In this work, we show that a simple method based on attribution patching outperforms all existing methods while requiring just two forward passes and a backward pass. We apply a linear approximation to activation patching to estimate the importance of each edge in the computational subgraph. Using this approximation, we prune the least important edges of the network. We survey the performance and limitations of this method, finding that averaged over all tasks our method has greater AUC from circuit recovery than other methods.
✨ ログイン状態を確認しています… PDF 被引用 BibTeX を表示 BibTeX を閉じる BibTeX を表示 引用
Causal Mediation Analysis for… Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias An overview of 11 proposals for building… An overview of 11 proposals for building safe advanced AI Causal Abstractions of Neural Networks Causal Abstractions of Neural Networks Low-Complexity Probing via Finding Subnetworks Low-Complexity Probing via Finding Subnetworks How does GPT-2 compute greater-than?… How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model Does Circuit Analysis Interpretability Scale?… Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla A Survey on Model Compression for Large… A Survey on Model Compression for Large Language Models Towards Automated Circuit Discovery for… Towards Automated Circuit Discovery for Mechanistic Interpretability Have Faith in Faithfulness: Going… Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms Sparse Autoencoders Enable Scalable and… Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models A Practical Review of Mechanistic… A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models A Primer on the Inner Workings of… A Primer on the Inner Workings of Transformer-based Language Models InversionView: A General-Purpose Method… InversionView: A General-Purpose Method for Reading Information from Neural Activations Decomposing and Editing Predictions by Modeling… Decomposing and Editing Predictions by Modeling Model Computation Unboxing the Black Box: Mechanistic… Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks Efficient Automated Circuit Discovery in… Efficient Automated Circuit Discovery in Transformers using Contextual Decomposition EAP-GP: Mitigating Saturation Effect in… EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification Interpreting vision transformers via… Interpreting vision transformers via residual replacement model Discovering Transformer Circuits via a Hybrid… Discovering Transformer Circuits via a Hybrid Attribution and Pruning Framework Attribution Patching Outperforms Automated… Attribution Patching Outperforms Automated Circuit Discovery 過去の参考文献 中心の論文 この論文を引用する論文 古い 新しい ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。