Dissecting Recall of Factual Associations in Auto-Regressive Language Models

Transformer-based language models (LMs) are known to capture factual knowledge in their parameters. While previous work looked into where factual associations are stored, only little is known about how they are retrieved internally during inference. We investigate this question through the lens of information flow. Given a subject-relation query, we study how the model aggregates information about the subject and relation to predict the correct attribute. With interventions on attention edges, we first identify two critical points where information propagates to the prediction: one from the relation positions followed by another from the subject positions. Next, by analyzing the information at these points, we unveil a three-step internal mechanism for attribute extraction. First, the representation at the last-subject position goes through an enrichment process, driven by the early MLP sublayers, to encode many subject-related attributes. Second, information from the relation propagates to the prediction. Third, the prediction representation "queries" the enriched subject to extract the attribute. Perhaps surprisingly, this extraction is typically done via attention heads, which often encode subject-attribute mappings in their parameters. Overall, our findings introduce a comprehensive view of how factual associations are stored and extracted internally in LMs, facilitating future research on knowledge localization and editing.

Layer NormalizationLayer NormalizationVisualizing andMeasuring the Geometry…Visualizing and Measuring the Geometry of BERTBERTnesia: Investigatingthe capture and…BERTnesia: Investigating the capture and forgetting of knowledge in BERTLocating and EditingFactual Associations in…Locating and Editing Factual Associations in GPTKnowledge Neurons inPretrained TransformersKnowledge Neurons in Pretrained TransformersAnalyzing Transformersin Embedding SpaceAnalyzing Transformers in Embedding SpaceInterpretability in theWild: a Circuit for…Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallMass-Editing Memory in aTransformerMass-Editing Memory in a TransformerProgress measures forgrokking via mechanisti…Progress measures for grokking via mechanistic interpretabilityMeasuring andManipulating Knowledge…Measuring and Manipulating Knowledge Representations in Language ModelsDoes Localization InformEditing? Surprising…Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language ModelsNeuromodulatory ControlNetworks (NCNs): A…Neuromodulatory Control Networks (NCNs): A Biologically Inspired Architecture for Dynamic LLM ProcessingA MechanisticInterpretation of…A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation AnalysisDoes Circuit AnalysisInterpretability Scale?…Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in ChinchillaTowards Best Practicesof Activation Patching…Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsPMET: Precise ModelEditing in a TransformerPMET: Precise Model Editing in a TransformerSumming Up the Facts:Additive Mechanisms…Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMsFine-Tuning EnhancesExisting Mechanisms: A…Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity TrackingA MechanisticUnderstanding of…A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and ToxicityIs This the Subspace YouAre Looking for? An…Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation PatchingInterpreting ArithmeticMechanism in Large…Interpreting Arithmetic Mechanism in Large Language Models through Comparative Neuron AnalysisAttention Satisfies: AConstraint-Satisfaction…Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language ModelsLanguage ModelsRepresent Space and TimeLanguage Models Represent Space and TimeDo I Know This Entity?Knowledge Awareness and…Do I Know This Entity? Knowledge Awareness and Hallucinations in Language ModelsDissecting Recall ofFactual Associations in…Dissecting Recall of Factual Associations in Auto-Regressive Language ModelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.