Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning

In-context learning (ICL) emerges as a promising capability of large language models (LLMs) by providing them with demonstration examples to perform diverse tasks. However, the underlying mechanism of how LLMs learn from the provided context remains under-explored. In this paper, we investigate the working mechanism of ICL through an information flow lens. Our findings reveal that label words in the demonstration examples function as anchors: (1) semantic information aggregates into label word representations during the shallow computation layers' processing; (2) the consolidated information in label words serves as a reference for LLMs' final predictions. Based on these insights, we introduce an anchor re-weighting method to improve ICL performance, a demonstration compression technique to expedite inference, and an analysis framework for diagnosing ICL errors in GPT2-XL. The promising applications of our findings again validate the uncovered ICL working mechanism and pave the way for future studies.

Recursive Deep Modelsfor Semantic…Recursive Deep Models for Semantic Compositionality Over a Sentiment TreebankAdam: A Method forStochastic OptimizationAdam: A Method for Stochastic OptimizationCharacter-levelConvolutional Networks…Character-level Convolutional Networks for Text ClassificationChain of ThoughtPrompting Elicits…Chain of Thought Prompting Elicits Reasoning in Large Language ModelsNoisy Channel LanguageModel Prompting for…Noisy Channel Language Model Prompting for Few-Shot Text ClassificationWhat Makes GoodIn-Context Examples for…What Makes Good In-Context Examples for GPT-3?Why Can GPT LearnIn-Context? Language…Why Can GPT Learn In-Context? Language Models Secretly Perform Gradient Descent as Meta-OptimizersUnified DemonstrationRetriever for In-Contex…Unified Demonstration Retriever for In-Context LearningTransformers asAlgorithms…Transformers as Algorithms: Generalization and Implicit Model Selection in In-context LearningLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsEvaluating theEffectiveness of Large…Evaluating the Effectiveness of Large Language Models in Representing Textual Descriptions of Geometry and Spatial Relations (Short Paper)What learning algorithmis in-context learning?…What learning algorithm is in-context learning? Investigations with linear modelsHow do Large LanguageModels Learn In-Context…How do Large Language Models Learn In-Context? Query and Key Matrices of In-Context Heads are Two Towers for Metric LearningUnveiling In-ContextLearning: A Coordinate…Unveiling In-Context Learning: A Coordinate System to Understand Its Working MechanismFunction Vectors inLarge Language ModelsFunction Vectors in Large Language ModelsPyramidKV: Dynamic KVCache Compression based…PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information FunnelingNeuron-Level KnowledgeAttribution in Large…Neuron-Level Knowledge Attribution in Large Language ModelsEmergent Abilities inLarge Language Models…Emergent Abilities in Large Language Models: A SurveyFrom Redundancy toRelevance: Enhancing…From Redundancy to Relevance: Enhancing Explainability in Multimodal Large Language ModelsUnderstanding theLanguage Model to Solve…Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer MechanismImplicit In-contextLearningImplicit In-context LearningRevisiting In-contextLearning Inference…Revisiting In-context Learning Inference Circuit in Large Language ModelsM2IV: Towards Efficientand Fine-grained…M2IV: Towards Efficient and Fine-grained Multimodal In-Context Learning in Large Vision-Language ModelsAnchor function: a typeof benchmark functions…Anchor function: a type of benchmark functions for studying language modelsLabel Words are Anchors:An Information Flow…Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。