Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion

This paper presents Diffusion Forcing, a new training paradigm where a diffusion model is trained to denoise a set of tokens with independent per-token noise levels. We apply Diffusion Forcing to sequence generative modeling by training a causal next-token prediction model to generate one or several future tokens without fully diffusing past ones. Our approach is shown to combine the strengths of next-token prediction models, such as variable-length generation, with the strengths of full-sequence diffusion models, such as the ability to guide sampling to desirable trajectories. Our method offers a range of additional capabilities, such as (1) rolling-out sequences of continuous tokens, such as video, with lengths past the training horizon, where baselines diverge and (2) new sampling and guiding schemes that uniquely profit from Diffusion Forcing's variable-horizon and causal architecture, and which lead to marked performance gains in decision-making and planning tasks. In addition to its empirical success, our method is proven to optimize a variational lower bound on the likelihoods of all subsequences of tokens drawn from the true joint distribution. Project website: https://boyuan.space/diffusion-forcing

Model Predictive PathIntegral Control using…Model Predictive Path Integral Control using Covariance Variable Importance SamplingGAIA-1: A GenerativeWorld Model for…GAIA-1: A Generative World Model for Autonomous DrivingIs ConditionalGenerative Modeling all…Is Conditional Generative Modeling all you need for Decision Making?Rolling Diffusion ModelsRolling Diffusion ModelsTutorial on DiffusionModels for Imaging and…Tutorial on Diffusion Models for Imaging and VisionPlayable Game GenerationPlayable Game GenerationDiffusion Models AreReal-Time Game EnginesDiffusion Models Are Real-Time Game EnginesContext as Memory:Scene-Consistent…Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory RetrievalA Survey of InteractiveGenerative VideoA Survey of Interactive Generative VideoFrom Slow Bidirectionalto Fast Autoregressive…From Slow Bidirectional to Fast Autoregressive Video Diffusion ModelsPosition: InteractiveGenerative Video as…Position: Interactive Generative Video as Next-Generation Game EngineDiffusion Policy PolicyOptimizationDiffusion Policy Policy OptimizationWorldPack: CompressedMemory Improves Spatial…WorldPack: Compressed Memory Improves Spatial Consistency in Video World ModelingError Analyses ofAuto-Regressive Video…Error Analyses of Auto-Regressive Video Diffusion Models: A Unified FrameworkEvaluating RobotPolicies in a World…Evaluating Robot Policies in a World ModelTaming Teacher Forcingfor Masked…Taming Teacher Forcing for Masked Autoregressive Video GenerationUnified World Models:Coupling Video and…Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic DatasetsDiffusion Forcing:Next-token Prediction…Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.