How to use and interpret activation patching

Activation patching is a popular mechanistic interpretability technique, but has many subtleties regarding how it is applied and how one may interpret the results. We provide a summary of advice and best practices, based on our experience using this technique in practice. We include an overview of the different ways to apply activation patching and a discussion on how to interpret the results. We focus on what evidence patching experiments provide about circuits, and on the choice of metric and associated pitfalls.

Causal Abstractions ofNeural NetworksCausal Abstractions of Neural NetworksCausal Analysis ofSyntactic Agreement…Causal Analysis of Syntactic Agreement Mechanisms in Neural Language ModelsLocating and EditingFactual Associations in…Locating and Editing Factual Associations in GPTTowards AutomatedCircuit Discovery for…Towards Automated Circuit Discovery for Mechanistic InterpretabilityDoes Circuit AnalysisInterpretability Scale?…Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in ChinchillaThe Hydra Effect:Emergent Self-repair in…The Hydra Effect: Emergent Self-repair in Language Model ComputationsInterpretability in theWild: a Circuit for…Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallLocalizing ModelBehavior with Path…Localizing Model Behavior with Path PatchingDissecting Recall ofFactual Associations in…Dissecting Recall of Factual Associations in Auto-Regressive Language ModelsCopy Suppression:Comprehensively…Copy Suppression: Comprehensively Understanding an Attention HeadA MechanisticInterpretation of…A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation AnalysisTowards Best Practicesof Activation Patching…Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsTranscoders FindInterpretable LLM…Transcoders Find Interpretable LLM Feature CircuitsA Practical Review ofMechanistic…A Practical Review of Mechanistic Interpretability for Transformer-Based Language ModelsA Primer on the InnerWorkings of…A Primer on the Inner Workings of Transformer-based Language ModelsMissed Causes andAmbiguous Effects…Missed Causes and Ambiguous Effects: Counterfactuals Pose Challenges for Interpreting Neural NetworksAttention Heads of LargeLanguage Models: A…Attention Heads of Large Language Models: A SurveyAddressing divergentrepresentations from…Addressing divergent representations from causal interventions on neural networksUnboxing the Black Box:Mechanistic…Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural NetworksHow do Large LanguageModels Understand…How do Large Language Models Understand Relevance? A Mechanistic Interpretability PerspectiveDo I Know This Entity?Knowledge Awareness and…Do I Know This Entity? Knowledge Awareness and Hallucinations in Language ModelsAligned but Blind:Alignment Increases…Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of RaceRepresentationEngineering for…Representation Engineering for Large-Language Models: Survey and Research ChallengesRethinking CircuitCompleteness in Languag…Rethinking Circuit Completeness in Language Models: AND, OR, and ADDER GatesHow to use and interpretactivation patchingHow to use and interpret activation patching過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。