Finding Neurons in a Haystack: Case Studies with Sparse Probing

Despite rapid adoption and deployment of large language models (LLMs), the internal computations of these models remain opaque and poorly understood. In this work, we seek to understand how high-level human-interpretable features are represented within the internal neuron activations of LLMs. We train $k$-sparse linear classifiers (probes) on these internal activations to predict the presence of features in the input; by varying the value of $k$ we study the sparsity of learned representations and how this varies with model scale. With $k=1$, we localize individual neurons which are highly relevant for a particular feature, and perform a number of case studies to illustrate general properties of LLMs. In particular, we show that early layers make use of sparse combinations of neurons to represent many features in superposition, that middle layers have seemingly dedicated neurons to represent higher-level contextual features, and that increasing scale causes representational sparsity to increase on average, but there are multiple types of scaling dynamics. In all, we probe for over 100 unique features comprising 10 different categories in 7 different models spanning 70 million to 6.9 billion parameters.

An InterpretabilityIllusion for BERTAn Interpretability Illusion for BERTTransformer Feed-ForwardLayers Are Key-Value…Transformer Feed-Forward Layers Are Key-Value MemoriesToy Models ofSuperpositionToy Models of SuperpositionTransformer Feed-ForwardLayers Build Prediction…Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpacePolysemanticity andCapacity in Neural…Polysemanticity and Capacity in Neural NetworksLocalizing ModelBehavior with Path…Localizing Model Behavior with Path PatchingDoes Localization InformEditing? Surprising…Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language ModelsInterpretability in theWild: a Circuit for…Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallAnalyzing Transformersin Embedding SpaceAnalyzing Transformers in Embedding SpaceEmergent WorldRepresentations…Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic TaskThe Quantization Modelof Neural ScalingThe Quantization Model of Neural ScalingProgress measures forgrokking via mechanisti…Progress measures for grokking via mechanistic interpretabilityEmergent LinearRepresentations in Worl…Emergent Linear Representations in World Models of Self-Supervised Sequence ModelsA MechanisticInterpretation of…A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation AnalysisDoes Circuit AnalysisInterpretability Scale?…Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in ChinchillaUniversal Neurons inGPT2 Language ModelsUniversal Neurons in GPT2 Language ModelsLanguage ModelsRepresent Space and TimeLanguage Models Represent Space and TimeTowards Best Practicesof Activation Patching…Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsTranscoders FindInterpretable LLM…Transcoders Find Interpretable LLM Feature CircuitsOpening the Black Box ofLarge Language Models…Opening the Black Box of Large Language Models: Two Views on Holistic InterpretabilityNeuron-Level KnowledgeAttribution in Large…Neuron-Level Knowledge Attribution in Large Language ModelsIs This the Subspace YouAre Looking for? An…Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation PatchingTowards PrincipledEvaluations of Sparse…Towards Principled Evaluations of Sparse Autoencoders for Interpretability and ControlScaling and evaluatingsparse autoencodersScaling and evaluating sparse autoencodersFinding Neurons in aHaystack: Case Studies…Finding Neurons in a Haystack: Case Studies with Sparse ProbingEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.