Universal Neurons in GPT2 Language Models

A basic question within the emerging field of mechanistic interpretability is the degree to which neural networks learn the same underlying mechanisms. In other words, are neural mechanisms universal across different models? In this work, we study the universality of individual neurons across GPT2 models trained from different initial random seeds, motivated by the hypothesis that universal neurons are likely to be interpretable. In particular, we compute pairwise correlations of neuron activations over 100 million tokens for every neuron pair across five different seeds and find that 1-5\% of neurons are universal, that is, pairs of neurons which consistently activate on the same inputs. We then study these universal neurons in detail, finding that they usually have clear interpretations and taxonomize them into a small number of neuron families. We conclude by studying patterns in neuron weights to establish several universal functional roles of neurons in simple circuits: deactivating attention heads, changing the entropy of the next token distribution, and predicting the next token to (not) be within a particular set.

Residual ConnectionsEncourage Iterative…Residual Connections Encourage Iterative InferenceIdentifying andControlling Important…Identifying and Controlling Important Neurons in Neural Machine TranslationFinding Neurons in aHaystack: Case Studies…Finding Neurons in a Haystack: Case Studies with Sparse ProbingTowards AutomatedCircuit Discovery for…Towards Automated Circuit Discovery for Mechanistic InterpretabilityEliciting LatentPredictions from…Eliciting Latent Predictions from Transformers with the Tuned LensThe Hydra Effect:Emergent Self-repair in…The Hydra Effect: Emergent Self-repair in Language Model ComputationsCopy Suppression:Comprehensively…Copy Suppression: Comprehensively Understanding an Attention HeadLanguage ModelsRepresent Space and TimeLanguage Models Represent Space and TimeSparse Autoencoders FindHighly Interpretable…Sparse Autoencoders Find Highly Interpretable Features in Language ModelsFunction Vectors inLarge Language ModelsFunction Vectors in Large Language ModelsNeurons in LargeLanguage Models: Dead…Neurons in Large Language Models: Dead, N-gram, PositionalHow do Language ModelsBind Entities in…How do Language Models Bind Entities in Context?The RemarkableRobustness of LLMs…The Remarkable Robustness of LLMs: Stages of Inference?A Practical Review ofMechanistic…A Practical Review of Mechanistic Interpretability for Transformer-Based Language ModelsLanguage-SpecificNeurons: The Key to…Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language ModelsUniversal Response andEmergence of Induction…Universal Response and Emergence of Induction in LLMsAn Investigation ofNeuron Activation as a…An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMsSparse AutoencodersReveal Universal Featur…Sparse Autoencoders Reveal Universal Feature Spaces Across Large Language ModelsActive-Dormant AttentionHeads: Mechanistically…Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMsAssessing theBrittleness of Safety…Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank ModificationsAutomaticallyInterpreting Millions o…Automatically Interpreting Millions of Features in Large Language ModelsTowards UnderstandingSafety Alignment: A…Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety NeuronsTranscoders Beat SparseAutoencoders for…Transcoders Beat Sparse Autoencoders for InterpretabilityWhen Models ManipulateManifolds: The Geometry…When Models Manipulate Manifolds: The Geometry of a Counting TaskUniversal Neurons inGPT2 Language ModelsUniversal Neurons in GPT2 Language ModelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.