Sparse Autoencoders Find Highly Interpretable Features in Language Models

One of the roadblocks to a better understanding of neural networks' internals is \textit{polysemanticity}, where neurons appear to activate in multiple, semantically distinct contexts. Polysemanticity prevents us from identifying concise, human-understandable explanations for what neural networks are doing internally. One hypothesised cause of polysemanticity is \textit{superposition}, where neural networks represent more features than they have neurons by assigning features to an overcomplete set of directions in activation space, rather than to individual neurons. Here, we attempt to identify those directions, using sparse autoencoders to reconstruct the internal activations of a language model. These autoencoders learn sets of sparsely activating features that are more interpretable and monosemantic than directions identified by alternative approaches, where interpretability is measured by automated methods. Moreover, we show that with our learned set of features, we can pinpoint the features that are causally responsible for counterfactual behaviour on the indirect object identification task \citep{wang2022interpretability} to a finer degree than previous decompositions. This work indicates that it is possible to resolve superposition in language models using a scalable, unsupervised method. Our method may serve as a foundation for future mechanistic interpretability work, which we hope will enable greater model transparency and steerability.

The Lottery TicketHypothesis: Finding…The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural NetworksAnalysis of theOptimization Landscapes…Analysis of the Optimization Landscapes for Overcomplete Representation LearningThe Pile: An 800GBDataset of Diverse Text…The Pile: An 800GB Dataset of Diverse Text for Language ModelingTransformervisualization via…Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factorsToy Models ofSuperpositionToy Models of SuperpositionLLM.int8(): 8-bit MatrixMultiplication for…LLM.int8(): 8-bit Matrix Multiplication for Transformers at ScaleInterpretability in theWild: a Circuit for…Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallTowards AutomatedCircuit Discovery for…Towards Automated Circuit Discovery for Mechanistic InterpretabilityAn Overview ofCatastrophic AI RisksAn Overview of Catastrophic AI RisksThe Alignment Problemfrom a Deep Learning…The Alignment Problem from a Deep Learning PerspectiveRefusal in LanguageModels Is Mediated by a…Refusal in Language Models Is Mediated by a Single DirectionLanguage ModelsRepresent Space and TimeLanguage Models Represent Space and TimeMathematical Models ofComputation in…Mathematical Models of Computation in SuperpositionWav-KAN: WaveletKolmogorov-Arnold…Wav-KAN: Wavelet Kolmogorov-Arnold NetworksDictionary LearningImproves Patch-Free…Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPTScaling and evaluatingsparse autoencodersScaling and evaluating sparse autoencodersLearning Multi-LevelFeatures with Matryoshk…Learning Multi-Level Features with Matryoshka Sparse AutoencodersAutomaticallyInterpreting Millions o…Automatically Interpreting Millions of Features in Large Language ModelsThe Geometry of Refusalin Large Language…The Geometry of Refusal in Large Language Models: Concept Cones and Representational IndependenceActivation SpaceInterventions Can Be…Activation Space Interventions Can Be Transferred Between Large Language ModelsUniversal SparseAutoencoders…Universal Sparse Autoencoders: Interpretable Cross-Model Concept AlignmentFrom Flat toHierarchical: Extractin…From Flat to Hierarchical: Extracting Sparse Representations with Matching PursuitSparse Autoencoders FindHighly Interpretable…Sparse Autoencoders Find Highly Interpretable Features in Language ModelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.