Towards Automated Circuit Discovery for Mechanistic Interpretability

Through considerable effort and intuition, several recent works have reverse-engineered nontrivial behaviors of transformer models. This paper systematizes the mechanistic interpretability process they followed. First, researchers choose a metric and dataset that elicit the desired model behavior. Then, they apply activation patching to find which abstract neural network units are involved in the behavior. By varying the dataset, metric, and units under investigation, researchers can understand the functionality of each component. We automate one of the process' steps: to identify the circuit that implements the specified behavior in the model's computational graph. We propose several algorithms and reproduce previous interpretability results to validate them. For example, the ACDC algorithm rediscovered 5/5 of the component types in a circuit in GPT-2 Small that computes the Greater-Than operation. ACDC selected 68 of the 32,000 edges in GPT-2 Small, all of which were manually found by previous work. Our code is available at https://github.com/ArthurConmy/Automatic-Circuit-Discovery.

An introduction to ROCanalysisAn introduction to ROC analysisBERT: Pre-training ofDeep Bidirectional…BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingScaling Laws for NeuralLanguage ModelsScaling Laws for Neural Language ModelsCausal Abstractions ofNeural NetworksCausal Abstractions of Neural NetworksEmergent Abilities ofLarge Language ModelsEmergent Abilities of Large Language ModelsHow does GPT-2 computegreater-than?…How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language modelLocalizing ModelBehavior with Path…Localizing Model Behavior with Path PatchingInterpretability atScale: Identifying…Interpretability at Scale: Identifying Causal Mechanisms in AlpacaFinding Neurons in aHaystack: Case Studies…Finding Neurons in a Haystack: Case Studies with Sparse ProbingCopy Suppression:Comprehensively…Copy Suppression: Comprehensively Understanding an Attention HeadGPT-4 Technical ReportGPT-4 Technical ReportAttribution PatchingOutperforms Automated…Attribution Patching Outperforms Automated Circuit DiscoveryThe Hydra Effect:Emergent Self-repair in…The Hydra Effect: Emergent Self-repair in Language Model ComputationsDoes Circuit AnalysisInterpretability Scale?…Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in ChinchillaThe Clock and the Pizza:Two Stories in…The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural NetworksRigorously AssessingNatural Language…Rigorously Assessing Natural Language Explanations of NeuronsTranscoders FindInterpretable LLM…Transcoders Find Interpretable LLM Feature CircuitsHow to use and interpretactivation patchingHow to use and interpret activation patchingIs This the Subspace YouAre Looking for? An…Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation PatchingUniversal Neurons inGPT2 Language ModelsUniversal Neurons in GPT2 Language ModelsSparse Autoencoders FindHighly Interpretable…Sparse Autoencoders Find Highly Interpretable Features in Language ModelsSumming Up the Facts:Additive Mechanisms…Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMsDictionary LearningImproves Patch-Free…Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPTAutomaticallyIdentifying Local and…Automatically Identifying Local and Global Circuits with Linear Computation GraphsTowards AutomatedCircuit Discovery for…Towards Automated Circuit Discovery for Mechanistic Interpretability過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。