A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis

Mathematical reasoning in large language models (LMs) has garnered significant attention in recent work, but there is a limited understanding of how these models process and store information related to arithmetic tasks within their architecture. In order to improve our understanding of this aspect of language models, we present a mechanistic interpretation of Transformer-based LMs on arithmetic questions using a causal mediation analysis framework. By intervening on the activations of specific model components and measuring the resulting changes in predicted probabilities, we identify the subset of parameters responsible for specific predictions. This provides insights into how information related to arithmetic is processed by LMs. Our experimental results indicate that LMs process the input by transmitting the information relevant to the query from mid-sequence early layers to the final token using the attention mechanism. Then, this information is processed by a set of MLP modules, which generate result-related information that is incorporated into the residual stream. To assess the specificity of the observed activation dynamics, we compare the effects of different model components on arithmetic queries with other tasks, including number retrieval from prompts and factual knowledge questions.

Large Language Modelsare Zero-Shot ReasonersLarge Language Models are Zero-Shot ReasonersImpact of PretrainingTerm Frequencies on…Impact of Pretraining Term Frequencies on Few-Shot Numerical ReasoningChain of ThoughtPrompting Elicits…Chain of Thought Prompting Elicits Reasoning in Large Language ModelsInterpretability in theWild: a Circuit for…Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallDissecting Recall ofFactual Associations in…Dissecting Recall of Factual Associations in Auto-Regressive Language ModelsFinding Neurons in aHaystack: Case Studies…Finding Neurons in a Haystack: Case Studies with Sparse ProbingProgress measures forgrokking via mechanisti…Progress measures for grokking via mechanistic interpretabilityLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsGoat: Fine-tuned LLaMAOutperforms GPT-4 on…Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic TasksSparks of ArtificialGeneral Intelligence…Sparks of Artificial General Intelligence: Early experiments with GPT-4Pythia: A Suite forAnalyzing Large Languag…Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingInference-TimeIntervention: Eliciting…Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelTowards a MechanisticInterpretation of…Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language ModelsTowards Best Practicesof Activation Patching…Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsPatchscopes: A UnifyingFramework for Inspectin…Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsA Practical Review ofMechanistic…A Practical Review of Mechanistic Interpretability for Transformer-Based Language ModelsHow to use and interpretactivation patchingHow to use and interpret activation patchingUnderstanding MultimodalLLMs: the Mechanistic…Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question AnsweringIs This the Subspace YouAre Looking for? An…Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation PatchingHave Faith inFaithfulness: Going…Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model MechanismsLanguage Models UseTrigonometry to Do…Language Models Use Trigonometry to Do AdditionBack Attention:Understanding and…Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language ModelsLarge Language Modelsand Causal Inference in…Large Language Models and Causal Inference in Collaboration: A Comprehensive SurveyWhen Models ManipulateManifolds: The Geometry…When Models Manipulate Manifolds: The Geometry of a Counting TaskA MechanisticInterpretation of…A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。