PMET: Precise Model Editing in a Transformer

Model editing techniques modify a minor proportion of knowledge in Large Language Models (LLMs) at a relatively low cost, which have demonstrated notable success. Existing methods assume Transformer Layer (TL) hidden states are values of key-value memories of the Feed-Forward Network (FFN). They usually optimize the TL hidden states to memorize target knowledge and use it to update the weights of the FFN in LLMs. However, the information flow of TL hidden states comes from three parts: Multi-Head Self-Attention (MHSA), FFN, and residual connections. Existing methods neglect the fact that the TL hidden states contains information not specifically required for FFN. Consequently, the performance of model editing decreases. To achieve more precise model editing, we analyze hidden states of MHSA and FFN, finding that MHSA encodes certain general knowledge extraction patterns. This implies that MHSA weights do not require updating when new knowledge is introduced. Based on above findings, we introduce PMET, which simultaneously optimizes Transformer Component (TC, namely MHSA and FFN) hidden states, while only using the optimized TC hidden states of FFN to precisely update FFN weights. Our experiments demonstrate that PMET exhibits state-of-the-art performance on both the \textsc{counterfact} and zsRE datasets. Our ablation experiments substantiate the effectiveness of our enhancements, further reinforcing the finding that the MHSA encodes certain general knowledge extraction patterns and indicating its storage of a small amount of factual knowledge. Our code is available at \url{https://github.com/xpq-tech/PMET}.

Zero-Shot RelationExtraction via Reading…Zero-Shot Relation Extraction via Reading ComprehensionEditable Neural NetworksEditable Neural NetworksModifying Memories inTransformer ModelsModifying Memories in Transformer ModelsEditing FactualKnowledge in Language…Editing Factual Knowledge in Language ModelsMemory-Based ModelEditing at ScaleMemory-Based Model Editing at ScaleLocating and EditingFactual Associations in…Locating and Editing Factual Associations in GPTEditing Large LanguageModels: Problems…Editing Large Language Models: Problems, Methods, and OpportunitiesCan We Edit FactualKnowledge by In-Context…Can We Edit Factual Knowledge by In-Context Learning?Mass-Editing Memory in aTransformerMass-Editing Memory in a TransformerDissecting Recall ofFactual Associations in…Dissecting Recall of Factual Associations in Auto-Regressive Language ModelsMeasuring andManipulating Knowledge…Measuring and Manipulating Knowledge Representations in Language ModelsEvaluating the RippleEffects of Knowledge…Evaluating the Ripple Effects of Knowledge Editing in Language ModelsEditing Large LanguageModels: Problems…Editing Large Language Models: Problems, Methods, and OpportunitiesCan We Edit MultimodalLarge Language Models?Can We Edit Multimodal Large Language Models?Memory Injections:Correcting Multi-Hop…Memory Injections: Correcting Multi-Hop Reasoning Failures During Inference in Transformer-Based Language ModelsUnveiling the Pitfallsof Knowledge Editing fo…Unveiling the Pitfalls of Knowledge Editing for Large Language ModelsModel Editing at Scaleleads to Gradual and…Model Editing at Scale leads to Gradual and Catastrophic ForgettingShould We Really EditLanguage Models? On the…Should We Really Edit Language Models? On the Evaluation of Edited Language ModelsEditable Fairness:Fine-Grained Bias…Editable Fairness: Fine-Grained Bias Mitigation in Language ModelsEditing KnowledgeRepresentation of…Editing Knowledge Representation of Language Lodel via Rephrased Prefix PromptsWilKE: Wise-LayerKnowledge Editor for…WilKE: Wise-Layer Knowledge Editor for Lifelong Knowledge EditingO-Edit: OrthogonalSubspace Editing for…O-Edit: Orthogonal Subspace Editing for Language Model Sequential EditingEva-KELLM: A NewBenchmark for Evaluatin…Eva-KELLM: A New Benchmark for Evaluating Knowledge Editing of LLMsEverything is Editable:Extend Knowledge Editin…Everything is Editable: Extend Knowledge Editing to Unstructured Data in Large Language ModelsPMET: Precise ModelEditing in a TransformerPMET: Precise Model Editing in a Transformer過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。