Linear Representations of Sentiment in Large Language Models

Sentiment is a pervasive feature in natural language text, yet it is an open question how sentiment is represented within Large Language Models (LLMs). In this study, we reveal that across a range of models, sentiment is represented linearly: a single direction in activation space mostly captures the feature across a range of tasks with one extreme for positive and the other for negative. Through causal interventions, we isolate this direction and show it is causally relevant in both toy tasks and real world datasets such as Stanford Sentiment Treebank. Through this case study we model a thorough investigation of what a single direction means on a broad data distribution. We further uncover the mechanisms that involve this direction, highlighting the roles of a small subset of attention heads and neurons. Finally, we discover a phenomenon which we term the summarization motif: sentiment is not solely represented on emotionally charged words, but is additionally summarized at intermediate positions without inherent sentiment, such as punctuation and names. We show that in Stanford Sentiment Treebank zero-shot classification, 76% of above-chance classification accuracy is lost when ablating the sentiment direction, nearly half of which (36%) is due to ablating the summarized sentiment direction exclusively at comma positions.

Open-DomainQuestion-AnsweringOpen-Domain Question-AnsweringSemEval-2016 Task 4:Sentiment Analysis in…SemEval-2016 Task 4: Sentiment Analysis in TwitterSemEval-2017 Task 4:Sentiment Analysis in…SemEval-2017 Task 4: Sentiment Analysis in TwitterBERT: Pre-training ofDeep Bidirectional…BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingNeural Natural LanguageInference Models…Neural Natural Language Inference Models Partially Embed Theories of Lexical Entailment and NegationDynaSent: A DynamicBenchmark for Sentiment…DynaSent: A Dynamic Benchmark for Sentiment AnalysisSurvey on sentimentanalysis: evolution of…Survey on sentiment analysis: evolution of research methods and topicsCausal Abstraction: ATheoretical Foundation…Causal Abstraction: A Theoretical Foundation for Mechanistic InterpretabilityTowards AutomatedCircuit Discovery for…Towards Automated Circuit Discovery for Mechanistic InterpretabilityOn the Origins of LinearRepresentations in Larg…On the Origins of Linear Representations in Large Language ModelsA Practical Review ofMechanistic…A Practical Review of Mechanistic Interpretability for Transformer-Based Language ModelsRefusal in LanguageModels Is Mediated by a…Refusal in Language Models Is Mediated by a Single DirectionHow to use and interpretactivation patchingHow to use and interpret activation patchingAttention Heads of LargeLanguage Models: A…Attention Heads of Large Language Models: A SurveyThe Geometry ofCategorical and…The Geometry of Categorical and Hierarchical Concepts in Large Language ModelsTowards PrincipledEvaluations of Sparse…Towards Principled Evaluations of Sparse Autoencoders for Interpretability and ControlCalibrating VerbalUncertainty as a Linear…Calibrating Verbal Uncertainty as a Linear Feature to Reduce HallucinationsRobust LLM safeguardingvia refusal feature…Robust LLM safeguarding via refusal feature adversarial trainingDo I Know This Entity?Knowledge Awareness and…Do I Know This Entity? Knowledge Awareness and Hallucinations in Language ModelsFrom Flat toHierarchical: Extractin…From Flat to Hierarchical: Extracting Sparse Representations with Matching PursuitLinear Representationsof Sentiment in Large…Linear Representations of Sentiment in Large Language ModelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.