Locating and Editing Factual Associations in GPT

We analyze the storage and recall of factual associations in autoregressive transformer language models, finding evidence that these associations correspond to localized, directly-editable computations. We first develop a causal intervention for identifying neuron activations that are decisive in a model's factual predictions. This reveals a distinct set of steps in middle-layer feed-forward modules that mediate factual predictions while processing subject tokens. To test our hypothesis that these computations correspond to factual association recall, we modify feed-forward weights to update specific factual associations using Rank-One Model Editing (ROME). We find that ROME is effective on a standard zero-shot relation extraction (zsRE) model-editing task, comparable to existing methods. To perform a more sensitive evaluation, we also evaluate ROME on a new dataset of counterfactual assertions, on which it simultaneously maintains both specificity and generalization, whereas other methods sacrifice one or another. Our results confirm an important role for mid-layer feed-forward modules in storing factual associations and suggest that direct manipulation of computational mechanisms may be a feasible approach for model editing. The code, dataset, visualizations, and an interactive demo notebook are available at https://rome.baulab.info/

Probing for semanticevidence of composition…Probing for semantic evidence of composition by means of simple classification tasksZero-Shot RelationExtraction via Reading…Zero-Shot Relation Extraction via Reading ComprehensionWhat you can cram into asingle $&!#* vector…What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic propertiesBERT: Pre-training ofDeep Bidirectional…BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingLanguage Models asKnowledge Bases?Language Models as Knowledge Bases?Modifying Memories inTransformer ModelsModifying Memories in Transformer ModelsHow Much Knowledge CanYou Pack Into the…How Much Knowledge Can You Pack Into the Parameters of a Language Model?BART: DenoisingSequence-to-Sequence…BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionFactual Probing Is[MASK]: Learning vs…Factual Probing Is [MASK]: Learning vs. Learning to RecallMeasuring and ImprovingConsistency in…Measuring and Improving Consistency in Pretrained Language ModelsDo Language Models HaveBeliefs? Methods for…Do Language Models Have Beliefs? Methods for Detecting, Updating, and Visualizing Model BeliefsMass-Editing Memory in aTransformerMass-Editing Memory in a TransformerExplaining Patterns inData with Language…Explaining Patterns in Data with Language Models via Interpretable AutopromptingEditing Large LanguageModels: Problems…Editing Large Language Models: Problems, Methods, and OpportunitiesHow Do Large LanguageModels Capture the…How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent AdvancesA MechanisticInterpretation of…A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation AnalysisPMET: Precise ModelEditing in a TransformerPMET: Precise Model Editing in a TransformerMELO: Enhancing ModelEditing with…MELO: Enhancing Model Editing with Neuron-Indexed Dynamic LoRAEditing FactualKnowledge and…Editing Factual Knowledge and Explanatory Ability of Medical Large Language ModelsSparse Feature Circuits:Discovering and Editing…Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsDo I Know This Entity?Knowledge Awareness and…Do I Know This Entity? Knowledge Awareness and Hallucinations in Language ModelsNAMET: Robust MassiveModel Editing via…NAMET: Robust Massive Model Editing via Noise-Aware Memory OptimizationPerturbation-RestrainedSequential Model EditingPerturbation-Restrained Sequential Model EditingTowardUltra-Long-Horizon…Toward Ultra-Long-Horizon Sequential Model EditingLocating and EditingFactual Associations in…Locating and Editing Factual Associations in GPTEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.