The Pitfalls of Next-Token Prediction

Can a mere next-token predictor faithfully model human intelligence? We crystallize this emerging concern and correct popular misconceptions surrounding it, and advocate a simple multi-token objective. As a starting point, we argue that the two often-conflated phases of next-token prediction -- autoregressive inference and teacher-forced training -- must be treated distinctly. The popular criticism that errors can compound during autoregressive inference, crucially assumes that teacher-forcing has learned an accurate next-token predictor. This assumption sidesteps a more deep-rooted problem we expose: in certain classes of tasks, teacher-forcing can simply fail to learn an accurate next-token predictor in the first place. We describe a general mechanism of how teacher-forcing can fail, and design a minimal planning task where both the Transformer and the Mamba architecture empirically fail in that manner -- remarkably, despite the task being straightforward to learn. Finally, we provide preliminary evidence that this failure can be resolved using _teacherless_ training, a simple modification using dummy tokens that predicts multiple tokens in advance. We hope this finding can ground future debates and inspire explorations beyond the next-token prediction paradigm. We make our code available under https://github.com/gregorbachmann/Next-Token-Failures

Google's Neural MachineTranslation System…Google's Neural Machine Translation System: Bridging the Gap between Human and Machine TranslationFine-Tuning LanguageModels from Human…Fine-Tuning Language Models from Human PreferencesTraining Verifiers toSolve Math Word ProblemsTraining Verifiers to Solve Math Word ProblemsShow Your Work:Scratchpads for…Show Your Work: Scratchpads for Intermediate Computation with Language ModelsIn-context Learning andInduction HeadsIn-context Learning and Induction HeadsLlama 2: Open Foundationand Fine-Tuned Chat…Llama 2: Open Foundation and Fine-Tuned Chat ModelsSparks of ArtificialGeneral Intelligence…Sparks of Artificial General Intelligence: Early experiments with GPT-4Embers ofAutoregression…Embers of Autoregression: Understanding Large Language Models Through the Problem They are Trained to SolvePositional DescriptionMatters for Transformer…Positional Description Matters for Transformers ArithmeticDistilling Step-by-Step!Outperforming Larger…Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model SizesAuto-RegressiveNext-Token Predictors…Auto-Regressive Next-Token Predictors are Universal LearnersTeaching Arithmetic toSmall TransformersTeaching Arithmetic to Small TransformersStream of Search (SoS):Learning to Search in…Stream of Search (SoS): Learning to Search in LanguageCausal Language ModelingCan Elicit Search and…Causal Language Modeling Can Elicit Search and Reasoning Capabilities on Logic PuzzlesIs Behavior Cloning AllYou Need? Understanding…Is Behavior Cloning All You Need? Understanding Horizon in Imitation LearningRetrieval-AugmentedGeneration with Graphs…Retrieval-Augmented Generation with Graphs (GraphRAG)Beyond Autoregression:Discrete Diffusion for…Beyond Autoregression: Discrete Diffusion for Complex Reasoning and PlanningDream 7B: DiffusionLarge Language ModelsDream 7B: Diffusion Large Language ModelsLaDiR: Latent DiffusionEnhances LLMs for Text…LaDiR: Latent Diffusion Enhances LLMs for Text ReasoningSystem 1.x: Learning toBalance Fast and Slow…System 1.x: Learning to Balance Fast and Slow Planning with Language ModelsComputational-StatisticalTradeoffs at the…Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier: Autoregressive and Imitation Learning under MisspecificationDream-VL & Dream-VLA:Open Vision-Language an…Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model BackboneLearning to Insert[PAUSE] Tokens for…Learning to Insert [PAUSE] Tokens for Better ReasoningOn Powerful Ways toGenerate…On Powerful Ways to Generate: Autoregression, Diffusion, and BeyondThe Pitfalls ofNext-Token PredictionThe Pitfalls of Next-Token PredictionEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.