Fast Model Editing at Scale

While large pre-trained models have enabled impressive results on a variety of downstream tasks, the largest existing models still make errors, and even accurate predictions may become outdated over time. Because detecting all such failures at training time is impossible, enabling both developers and end users of such models to correct inaccurate outputs while leaving the model otherwise intact is desirable. However, the distributed, black-box nature of the representations learned by large neural networks makes producing such targeted edits difficult. If presented with only a single problematic input and new desired output, fine-tuning approaches tend to overfit; other editing algorithms are either computationally infeasible or simply ineffective when applied to very large models. To enable easy post-hoc editing at scale, we propose Model Editor Networks using Gradient Decomposition (MEND), a collection of small auxiliary editing networks that use a single desired input-output pair to make fast, local edits to a pre-trained model's behavior. MEND learns to transform the gradient obtained by standard fine-tuning, using a low-rank decomposition of the gradient to make the parameterization of this transformation tractable. MEND can be trained on a single GPU in less than a day even for 10 billion+ parameter models; once trained MEND enables rapid application of new edits to the pre-trained model. Our experiments with T5, GPT, BERT, and BART models show that MEND is the only approach to model editing that effectively edits the behavior of models with more than 10 billion parameters. Code and data available at https://sites.google.com/view/mend-editing.

CatastrophicInterference in…Catastrophic Interference in Connectionist Networks: The Sequential Learning ProblemAdam: A Method forStochastic OptimizationAdam: A Method for Stochastic OptimizationDeep Residual Learningfor Image RecognitionDeep Residual Learning for Image RecognitionZero-Shot RelationExtraction via Reading…Zero-Shot Relation Extraction via Reading ComprehensionAttention Is All YouNeedAttention Is All You NeedOptimization as a Modelfor Few-Shot LearningOptimization as a Model for Few-Shot LearningMeta-SGD: Learning toLearn Quickly for Few…Meta-SGD: Learning to Learn Quickly for Few Shot LearningContinual LifelongLearning with Neural…Continual Lifelong Learning with Neural Networks: A ReviewHuggingFace'sTransformers…HuggingFace's Transformers: State-of-the-art Natural Language ProcessingGeneralized Inner LoopMeta-LearningGeneralized Inner Loop Meta-LearningEditing FactualKnowledge in Language…Editing Factual Knowledge in Language ModelsKnowledge Neurons inPretrained TransformersKnowledge Neurons in Pretrained TransformersMemory-Based ModelEditing at ScaleMemory-Based Model Editing at ScaleMemory-assisted promptediting to improve GPT-…Memory-assisted prompt editing to improve GPT-3 after deploymentVQGAN-CLIP: Open DomainImage Generation and…VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language GuidanceEditing Models with TaskArithmeticEditing Models with Task ArithmeticUnlearn What You Want toForget: Efficient…Unlearn What You Want to Forget: Efficient Unlearning for LLMsPost-hoc ConceptBottleneck ModelsPost-hoc Concept Bottleneck ModelsA Unified Framework forModel EditingA Unified Framework for Model EditingLarimar: Large LanguageModels with Episodic…Larimar: Large Language Models with Episodic Memory ControlWilKE: Wise-LayerKnowledge Editor for…WilKE: Wise-Layer Knowledge Editor for Lifelong Knowledge EditingCoME: AnUnlearning-based…CoME: An Unlearning-based Approach to Conflict-free Model EditingMitigating HeterogeneousToken Overfitting in LL…Mitigating Heterogeneous Token Overfitting in LLM Knowledge EditingDomain Specialization asthe Key to Make Large…Domain Specialization as the Key to Make Large Language Models Disruptive: A Comprehensive SurveyFast Model Editing atScaleFast Model Editing at ScaleEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.