In-Context Learning Creates Task Vectors

In-context learning (ICL) in Large Language Models (LLMs) has emerged as a powerful new learning paradigm. However, its underlying mechanism is still not well understood. In particular, it is challenging to map it to the "standard" machine learning framework, where one uses a training set $S$ to find a best-fitting function $f(x)$ in some hypothesis class. Here we make progress on this problem by showing that the functions learned by ICL often have a very simple structure: they correspond to the transformer LLM whose only inputs are the query $x$ and a single "task vector" calculated from the training set. Thus, ICL can be seen as compressing $S$ into a single task vector $\boldsymbolθ(S)$ and then using this task vector to modulate the transformer to produce the output. We support the above claim via comprehensive experiments across a range of models and tasks.

Why Can GPT LearnIn-Context? Language…Why Can GPT Learn In-Context? Language Models Secretly Perform Gradient Descent as Meta-OptimizersAn Explanation ofIn-context Learning as…An Explanation of In-context Learning as Implicit Bayesian InferenceData DistributionalProperties Drive…Data Distributional Properties Drive Emergent In-Context Learning in TransformersIn-context Learning andInduction HeadsIn-context Learning and Induction HeadsA Survey for In-contextLearningA Survey for In-context LearningWhy Can GPT LearnIn-Context? Language…Why Can GPT Learn In-Context? Language Models Secretly Perform Gradient Descent as Meta-OptimizersWhat learning algorithmis in-context learning?…What learning algorithm is in-context learning? Investigations with linear modelsTransformers LearnIn-Context by Gradient…Transformers Learn In-Context by Gradient DescentEditing Models with TaskArithmeticEditing Models with Task ArithmeticLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsEvaluating theEffectiveness of Large…Evaluating the Effectiveness of Large Language Models in Representing Textual Descriptions of Geometry and Spatial Relations (Short Paper)Language ModelsImplement Simple…Language Models Implement Simple Word2Vec-style Vector ArithmeticThe LinearRepresentation…The Linear Representation Hypothesis and the Geometry of Large Language ModelsFinding Visual TaskVectorsFinding Visual Task VectorsUnifying Attention Headsand Task Vectors via…Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context LearningImplicit In-contextLearningImplicit In-context LearningICLR: In-ContextLearning of…ICLR: In-Context Learning of RepresentationsTask Vectors, LearnedNot Extracted…Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic InsightTowards GeneralizableImplicit In-Context…Towards Generalizable Implicit In-Context Learning with Attention RoutingEverything EverywhereAll at Once: LLMs can…Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in SuperpositionM2IV: Towards Efficientand Fine-grained…M2IV: Towards Efficient and Fine-grained Multimodal In-Context Learning in Large Vision-Language ModelsRevisiting VerilogEval:A Year of Improvements…Revisiting VerilogEval: A Year of Improvements in Large-Language Models for Hardware Code GenerationThe Geometry ofPrompting: Unveiling…The Geometry of Prompting: Unveiling Distinct Mechanisms of Task Adaptation in Language ModelsDistributed Rule Vectorsis A Key Mechanism in…Distributed Rule Vectors is A Key Mechanism in Large Language Models' In-Context LearningIn-Context LearningCreates Task VectorsIn-Context Learning Creates Task Vectors過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。