The Linear Representation Hypothesis and the Geometry of Large Language Models

Informally, the 'linear representation hypothesis' is the idea that high-level concepts are represented linearly as directions in some representation space. In this paper, we address two closely related questions: What does "linear representation" actually mean? And, how do we make sense of geometric notions (e.g., cosine similarity or projection) in the representation space? To answer these, we use the language of counterfactuals to give two formalizations of "linear representation", one in the output (word) representation space, and one in the input (sentence) space. We then prove these connect to linear probing and model steering, respectively. To make sense of geometric notions, we use the formalization to identify a particular (non-Euclidean) inner product that respects language structure in a sense we make precise. Using this causal inner product, we show how to unify all notions of linear representation. In particular, this allows the construction of probes and steering vectors using counterfactual pairs. Experiments with LLaMA-2 demonstrate the existence of linear representations of concepts, the connection to interpretation and control, and the fundamental role of the choice of inner product.

Toy Models ofSuperpositionToy Models of SuperpositionEmergent LinearRepresentations in Worl…Emergent Linear Representations in World Models of Self-Supervised Sequence ModelsIn-Context LearningCreates Task VectorsIn-Context Learning Creates Task VectorsRepresentationEngineering: A Top-Down…Representation Engineering: A Top-Down Approach to AI TransparencyActivation Addition:Steering Language Model…Activation Addition: Steering Language Models Without OptimizationLlama 2: Open Foundationand Fine-Tuned Chat…Llama 2: Open Foundation and Fine-Tuned Chat ModelsGPT-4 Technical ReportGPT-4 Technical ReportLanguage ModelsRepresent Space and TimeLanguage Models Represent Space and TimeFunction Vectors inLarge Language ModelsFunction Vectors in Large Language ModelsLinearity of RelationDecoding in Transformer…Linearity of Relation Decoding in Transformer Language ModelsLanguage ModelsImplement Simple…Language Models Implement Simple Word2Vec-style Vector ArithmeticGemma: Open Models Basedon Gemini Research and…Gemma: Open Models Based on Gemini Research and TechnologyImprovingInstruction-Following i…Improving Instruction-Following in Language Models through Activation SteeringRefusal in LanguageModels Is Mediated by a…Refusal in Language Models Is Mediated by a Single DirectionReFT: RepresentationFinetuning for Language…ReFT: Representation Finetuning for Language Models3-in-1: 2D RotaryAdaptation for Efficien…3-in-1: 2D Rotary Adaptation for Efficient Finetuning, Efficient Batching and ComposabilityNot All Language ModelFeatures Are…Not All Language Model Features Are One-Dimensionally LinearAxBench: Steering LLMs?Even Simple Baselines…AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse AutoencodersICLR: In-ContextLearning of…ICLR: In-Context Learning of RepresentationsFrom Flat toHierarchical: Extractin…From Flat to Hierarchical: Extracting Sparse Representations with Matching PursuitThe Geometry of Refusalin Large Language…The Geometry of Refusal in Large Language Models: Concept Cones and Representational IndependenceActivation SpaceInterventions Can Be…Activation Space Interventions Can Be Transferred Between Large Language ModelsDo LLMs "know"internally when they…Do LLMs "know" internally when they follow instructions?FairSteer: InferenceTime Debiasing for LLMs…FairSteer: Inference Time Debiasing for LLMs with Dynamic Activation SteeringThe LinearRepresentation…The Linear Representation Hypothesis and the Geometry of Large Language Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。