Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language Models

Language models learn a great quantity of factual information during pretraining, and recent work localizes this information to specific model weights like mid-layer MLP weights. In this paper, we find that we can change how a fact is stored in a model by editing weights that are in a different location than where existing methods suggest that the fact is stored. This is surprising because we would expect that localizing facts to specific model parameters would tell us where to manipulate knowledge in models, and this assumption has motivated past work on model editing methods. Specifically, we show that localization conclusions from representation denoising (also known as Causal Tracing) do not provide any insight into which model MLP layer would be best to edit in order to override an existing stored fact with a new one. This finding raises questions about how past work relies on Causal Tracing to select which model layers to edit. Next, we consider several variants of the editing problem, including erasing and amplifying facts. For one of our editing problems, editing performance does relate to localization results from representation denoising, but we find that which layer we edit is a far better predictor of performance. Our results suggest, counterintuitively, that better mechanistic understanding of how pretrained language models work may not always translate to insights about how to best change their behavior. Our code is available at https://github.com/google/belief-localization

Efficient Estimation ofWord Representations in…Efficient Estimation of Word Representations in Vector SpaceVisualizing andUnderstanding…Visualizing and Understanding Convolutional NetworksModifying Memories inTransformer ModelsModifying Memories in Transformer ModelsUnderstanding the Roleof Individual Units in…Understanding the Role of Individual Units in a Deep Neural NetworkDo Language Models HaveBeliefs? Methods for…Do Language Models Have Beliefs? Methods for Detecting, Updating, and Visualizing Model BeliefsAn InterpretabilityIllusion for BERTAn Interpretability Illusion for BERTKnowledge Neurons inPretrained TransformersKnowledge Neurons in Pretrained TransformersMemory-Based ModelEditing at ScaleMemory-Based Model Editing at ScaleFinding Skill Neurons inPre-trained…Finding Skill Neurons in Pre-trained Transformer-based Language ModelsFast Model Editing atScaleFast Model Editing at ScaleTransformer Feed-ForwardLayers Build Prediction…Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceMass-Editing Memory in aTransformerMass-Editing Memory in a TransformerAging with GRACE:Lifelong Model Editing…Aging with GRACE: Lifelong Model Editing with Discrete Key-Value AdaptorsEditing Large LanguageModels: Problems…Editing Large Language Models: Problems, Methods, and OpportunitiesCan We Edit MultimodalLarge Language Models?Can We Edit Multimodal Large Language Models?Emergent LinearRepresentations in Worl…Emergent Linear Representations in World Models of Self-Supervised Sequence ModelsFinding Neurons in aHaystack: Case Studies…Finding Neurons in a Haystack: Case Studies with Sparse ProbingUnveiling the Pitfallsof Knowledge Editing fo…Unveiling the Pitfalls of Knowledge Editing for Large Language ModelsDecomposing and EditingPredictions by Modeling…Decomposing and Editing Predictions by Modeling Model ComputationLinearity of RelationDecoding in Transformer…Linearity of Relation Decoding in Transformer Language ModelsIs This the Subspace YouAre Looking for? An…Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation PatchingCan SensitiveInformation Be Deleted…Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction AttacksNavigating the DualFacets: A Comprehensive…Navigating the Dual Facets: A Comprehensive Evaluation of Sequential Memory Editing in Large Language ModelsHow new data permeatesLLM knowledge and how t…How new data permeates LLM knowledge and how to dilute itDoes Localization InformEditing? Surprising…Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。