In-context Learning and Induction Heads

"Induction heads" are attention heads that implement a simple algorithm to complete token sequences like [A][B] ... [A] -> [B]. In this work, we present preliminary and indirect evidence for a hypothesis that induction heads might constitute the mechanism for the majority of all "in-context learning" in large transformer models (i.e. decreasing loss at increasing token indices). We find that induction heads develop at precisely the same point as a sudden sharp increase in in-context learning ability, visible as a bump in the training loss. We present six complementary lines of evidence, arguing that induction heads may be the mechanistic source of general in-context learning in transformer models of any size. For small attention-only models, we present strong, causal evidence; for larger models with MLPs, we present correlational evidence.

Neural MachineTranslation by Jointly…Neural Machine Translation by Jointly Learning to Align and TranslateAnalyzing Multi-HeadSelf-Attention…Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be PrunedWhat Does BERT Look At?An Analysis of BERT's…What Does BERT Look At? An Analysis of BERT's AttentionA MultiscaleVisualization of…A Multiscale Visualization of Attention in the Transformer ModelSimilarity of NeuralNetwork Representations…Similarity of Neural Network Representations RevisitedLanguage Models areFew-Shot LearnersLanguage Models are Few-Shot LearnersScaling Laws for NeuralLanguage ModelsScaling Laws for Neural Language ModelsEvaluating LargeLanguage Models Trained…Evaluating Large Language Models Trained on CodeA General LanguageAssistant as a…A General Language Assistant as a Laboratory for AlignmentAn Explanation ofIn-context Learning as…An Explanation of In-context Learning as Implicit Bayesian InferenceRethinking the Role ofDemonstrations: What…Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Grokking: GeneralizationBeyond Overfitting on…Grokking: Generalization Beyond Overfitting on Small Algorithmic DatasetsAttentionViz: A GlobalView of Transformer…AttentionViz: A Global View of Transformer AttentionA MechanisticInterpretation of…A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation AnalysisToward Transparent AI: ASurvey on Interpreting…Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural NetworksUnveiling InductionHeads: Provable Trainin…Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in TransformersAttention Heads of LargeLanguage Models: A…Attention Heads of Large Language Models: A SurveyAttention with Markov: AFramework for Principle…Attention with Markov: A Framework for Principled Analysis of Transformers via Markov ChainsIs Mamba Capable ofIn-Context Learning?Is Mamba Capable of In-Context Learning?What Can TransformerLearn with Varying…What Can Transformer Learn with Varying Depth? Case Studies on Sequence Learning TasksSimple linear attentionlanguage models balance…Simple linear attention language models balance the recall-throughput tradeoffIn-Context LearningStrategies Emerge…In-Context Learning Strategies Emerge RationallySoftmax is not Enough(for Sharp Size…Softmax is not Enough (for Sharp Size Generalisation)In-Context LinearRegression Demystified…In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax AttentionIn-context Learning andInduction HeadsIn-context Learning and Induction Heads過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。