Function Vectors in Large Language Models

We report the presence of a simple neural mechanism that represents an input-output function as a vector within autoregressive transformer language models (LMs). Using causal mediation analysis on a diverse range of in-context-learning (ICL) tasks, we find that a small number attention heads transport a compact representation of the demonstrated task, which we call a function vector (FV). FVs are robust to changes in context, i.e., they trigger execution of the task on inputs such as zero-shot and natural text settings that do not resemble the ICL contexts from which they are collected. We test FVs across a range of tasks, models, and layers and find strong causal effects across settings in middle layers. We investigate the internal structure of FVs and find while that they often contain information that encodes the output space of the function, this information alone is not sufficient to reconstruct an FV. Finally, we test semantic vector composition in FVs, and find that to some extent they can be summed to create vectors that trigger new complex tasks. Our findings show that compact, causal internal vector representations of function abstractions can be explicitly extracted from LLMs. Our code and data are available at https://functions.baulab.info.

In-context Learning andInduction HeadsIn-context Learning and Induction HeadsActivation Addition:Steering Language Model…Activation Addition: Steering Language Models Without OptimizationDissecting Recall ofFactual Associations in…Dissecting Recall of Factual Associations in Auto-Regressive Language ModelsRepresentationEngineering: A Top-Down…Representation Engineering: A Top-Down Approach to AI TransparencyEliciting LatentPredictions from…Eliciting Latent Predictions from Transformers with the Tuned LensThe Learnability ofIn-Context LearningThe Learnability of In-Context LearningLabel Words are Anchors:An Information Flow…Label Words are Anchors: An Information Flow Perspective for Understanding In-Context LearningLlama 2: Open Foundationand Fine-Tuned Chat…Llama 2: Open Foundation and Fine-Tuned Chat ModelsIn-context Vectors:Making In Context…In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space SteeringLinearity of RelationDecoding in Transformer…Linearity of Relation Decoding in Transformer Language ModelsLanguage ModelsImplement Simple…Language Models Implement Simple Word2Vec-style Vector ArithmeticWhat and How doesIn-Context Learning…What and How does In-Context Learning Learn? Bayesian Model Averaging, Parameterization, and GeneralizationImproving ActivationSteering in Language…Improving Activation Steering in Language Models with Mean-CentringThe LinearRepresentation…The Linear Representation Hypothesis and the Geometry of Large Language ModelsA MechanisticUnderstanding of…A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and ToxicityUniversal Neurons inGPT2 Language ModelsUniversal Neurons in GPT2 Language ModelsAssessing theBrittleness of Safety…Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank ModificationsLocating and ExtractingRelational Concepts in…Locating and Extracting Relational Concepts in Large Language ModelsFinding Visual TaskVectorsFinding Visual Task VectorsOpening the Black Box ofLarge Language Models…Opening the Black Box of Large Language Models: Two Views on Holistic InterpretabilitySpectral Filters, DarkSignals, and Attention…Spectral Filters, Dark Signals, and Attention SinksICLR: In-ContextLearning of…ICLR: In-Context Learning of RepresentationsA Data GenerationPerspective to the…A Data Generation Perspective to the Mechanism of In-Context LearningDo LLMs "know"internally when they…Do LLMs "know" internally when they follow instructions?Function Vectors inLarge Language ModelsFunction Vectors in Large Language Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。