Knowledge Neurons in Pretrained Transformers

Large-scale pretrained language models are surprisingly good at recalling factual knowledge presented in the training corpus (Petroni et al., 2019; Jiang et al., 2020b).In this paper, we present preliminary studies on how factual knowledge is stored in pretrained Transformers by introducing the concept of knowledge neurons.Specifically, we examine the fill-in-the-blank cloze task for BERT.Given a relational fact, we propose a knowledge attribution method to identify the neurons that express the fact.We find that the activation of such knowledge neurons is positively correlated to the expression of their corresponding facts.In our case studies, we attempt to leverage knowledge neurons to edit (such as update, and erase) specific factual knowledge without fine-tuning.Our results shed light on understanding the storage of knowledge within pretrained Transformers.The code is available at https://github.com/Hunter-DDM/knowledge-neurons.

Visualizing andUnderstanding…Visualizing and Understanding Convolutional NetworksAxiomatic Attributionfor Deep NetworksAxiomatic Attribution for Deep NetworksLearning ImportantFeatures Through…Learning Important Features Through Propagating Activation DifferencesT-REx: A Large ScaleAlignment of Natural…T-REx: A Large Scale Alignment of Natural Language with Knowledge Base TriplesTransformer-BasedFeature Learning for…Transformer-Based Feature Learning for Algorithm Selection in Combinatorial OptimisationLanguage Models asKnowledge Bases?Language Models as Knowledge Bases?BERT: Pre-training ofDeep Bidirectional…BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingCross-lingual LanguageModel PretrainingCross-lingual Language Model PretrainingCOMET: CommonsenseTransformers for…COMET: Commonsense Transformers for Automatic Knowledge Graph ConstructionUnified Language ModelPre-training for Natura…Unified Language Model Pre-training for Natural Language Understanding and GenerationTransformer Feed-ForwardLayers Are Key-Value…Transformer Feed-Forward Layers Are Key-Value MemoriesMeasuring and ImprovingConsistency in…Measuring and Improving Consistency in Pretrained Language ModelsBERTnesia: Investigatingthe capture and…BERTnesia: Investigating the capture and forgetting of knowledge in BERTDistilling RelationEmbeddings from…Distilling Relation Embeddings from Pretrained Language ModelsTowards TracingKnowledge in Language…Towards Tracing Knowledge in Language Models Back to the Training DataEditing Large LanguageModels: Problems…Editing Large Language Models: Problems, Methods, and OpportunitiesUnderstandingTransformer Memorizatio…Understanding Transformer Memorization Recall Through IdiomsA Survey onHallucination in Large…A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open QuestionsKnowledge Editing forLarge Language Models…Knowledge Editing for Large Language Models: A SurveyIs This the Subspace YouAre Looking for? An…Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation PatchingTowards InterpretingLanguage Models: A Case…Towards Interpreting Language Models: A Case Study in Multi-Hop ReasoningFrom Tokens to Words: Onthe Inner Lexicon of…From Tokens to Words: On the Inner Lexicon of LLMsPerturbation-RestrainedSequential Model EditingPerturbation-Restrained Sequential Model EditingParamMute: SuppressingKnowledge-Critical FFNs…ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented GenerationKnowledge Neurons inPretrained TransformersKnowledge Neurons in Pretrained TransformersEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.