Emergent Linear Representations in World Models of Self-Supervised Sequence Models

How do sequence models represent their decision-making process? Prior work suggests that Othello-playing neural network learned nonlinear models of the board state (Li et al., 2023a). In this work, we provide evidence of a closely related linear representation of the board. In particular, we show that probing for "my colour" vs. "opponent's colour" may be a simple yet powerful way to interpret the model's internal state. This precise understanding of the internal representations allows us to control the model's behaviour with simple vector arithmetic. Linear representations enable significant interpretability progress, which we demonstrate with further exploration of how the world model is computed.

Efficient Estimation ofWord Representations in…Efficient Estimation of Word Representations in Vector SpaceToy Models ofSuperpositionToy Models of SuperpositionLocating and EditingFactual Associations in…Locating and Editing Factual Associations in GPTEmergent WorldRepresentations…Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic TaskFinding Neurons in aHaystack: Case Studies…Finding Neurons in a Haystack: Case Studies with Sparse ProbingThe Geometry of Truth:Emergent Linear…The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False DatasetsInference-TimeIntervention: Eliciting…Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelDoes Localization InformEditing? Surprising…Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language ModelsThe Hydra Effect:Emergent Self-repair in…The Hydra Effect: Emergent Self-repair in Language Model ComputationsDiscovering LatentKnowledge in Language…Discovering Latent Knowledge in Language Models Without SupervisionMass-Editing Memory in aTransformerMass-Editing Memory in a TransformerProgress measures forgrokking via mechanisti…Progress measures for grokking via mechanistic interpretabilityLanguage ModelsRepresent Space and TimeLanguage Models Represent Space and TimeOn the Origins of LinearRepresentations in Larg…On the Origins of Linear Representations in Large Language ModelsTowards Best Practicesof Activation Patching…Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsRefusal in LanguageModels Is Mediated by a…Refusal in Language Models Is Mediated by a Single DirectionLearning InterpretableConcepts: Unifying…Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation ModelsOpening the Black Box ofLarge Language Models…Opening the Black Box of Large Language Models: Two Views on Holistic InterpretabilityIs This the Subspace YouAre Looking for? An…Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation PatchingNot All Language ModelFeatures Are…Not All Language Model Features Are One-Dimensionally LinearICLR: In-ContextLearning of…ICLR: In-Context Learning of RepresentationsAxBench: Steering LLMs?Even Simple Baselines…AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse AutoencodersFrom Flat toHierarchical: Extractin…From Flat to Hierarchical: Extracting Sparse Representations with Matching PursuitTowards PrincipledEvaluations of Sparse…Towards Principled Evaluations of Sparse Autoencoders for Interpretability and ControlEmergent LinearRepresentations in Worl…Emergent Linear Representations in World Models of Self-Supervised Sequence Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。