Knowledge Neurons in Pretrained Transformers

Large-scale pretrained language models are surprisingly good at recalling factual knowledge presented in the training corpus. In this paper, we present preliminary studies on how factual knowledge is stored in pretrained Transformers by introducing the concept of knowledge neurons. Specifically, we examine the fill-in-the-blank cloze task for BERT. Given a relational fact, we propose a knowledge attribution method to identify the neurons that express the fact. We find that the activation of such knowledge neurons is positively correlated to the expression of their corresponding facts. In our case studies, we attempt to leverage knowledge neurons to edit (such as update, and erase) specific factual knowledge without fine-tuning. Our results shed light on understanding the storage of knowledge within pretrained Transformers. The code is available at https://github.com/Hunter-DDM/knowledge-neurons.

MQuAKE: AssessingKnowledge Editing in…MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop QuestionsDoes Localization InformEditing? Surprising…Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language ModelsUnlearn What You Want toForget: Efficient…Unlearn What You Want to Forget: Efficient Unlearning for LLMsPhysics of LanguageModels: Part 3.1…Physics of Language Models: Part 3.1, Knowledge Storage and ExtractionA Unified Framework forModel EditingA Unified Framework for Model EditingDeepEdit: KnowledgeEditing as Decoding wit…DeepEdit: Knowledge Editing as Decoding with ConstraintsLarimar: Large LanguageModels with Episodic…Larimar: Large Language Models with Episodic Memory ControlA Survey on MechanisticInterpretability for…A Survey on Mechanistic Interpretability for Multi-Modal Foundation ModelsFrom Redundancy toRelevance: Enhancing…From Redundancy to Relevance: Enhancing Explainability in Multimodal Large Language ModelsIn-Context Editing:Learning Knowledge from…In-Context Editing: Learning Knowledge from Self-Induced DistributionsThe UnreasonableIneffectiveness of the…The Unreasonable Ineffectiveness of the Deeper LayersMitigating HeterogeneousToken Overfitting in LL…Mitigating Heterogeneous Token Overfitting in LLM Knowledge EditingKnowledge Neurons inPretrained TransformersKnowledge Neurons in Pretrained TransformersFocus paperCiting papersOlderNewer

No earlier referenced papers in the catalog yet.

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.