Dream 7B: Diffusion Large Language Models

We introduce Dream 7B, the most powerful open diffusion large language model to date. Unlike autoregressive (AR) models that generate tokens sequentially, Dream 7B employs discrete diffusion modeling to refine sequences in parallel through iterative denoising. Our model consistently outperforms existing diffusion language models on general, mathematical, and coding tasks. Dream 7B demonstrates superior planning abilities and inference flexibility, including arbitrary-order generation, infilling capabilities, and tunable quality-speed trade-offs. These results are achieved through simple yet effective training techniques, including AR-based LLM initialization and context-adaptive token-level noise rescheduling. We release both Dream-Base and Dream-Instruct to facilitate further research in diffusion-based language modeling.

Training Verifiers toSolve Math Word ProblemsTraining Verifiers to Solve Math Word ProblemsContinuous diffusion forcategorical dataContinuous diffusion for categorical dataGPT-4 Technical ReportGPT-4 Technical ReportThe Llama 3 Herd ofModelsThe Llama 3 Herd of ModelsTÜLU 3: PushingFrontiers in Open…TÜLU 3: Pushing Frontiers in Open Language Model Post-TrainingScaling LLM Test-TimeCompute Optimally can b…Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model ParametersBlock Diffusion:Interpolating Between…Block Diffusion: Interpolating Between Autoregressive and Diffusion Language ModelsLarge Language DiffusionModelsLarge Language Diffusion ModelsScaling up MaskedDiffusion Models on TextScaling up Masked Diffusion Models on TextDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningSmolLM2: When Smol GoesBig - Data-Centric…SmolLM2: When Smol Goes Big - Data-Centric Training of a Small Language ModelScaling up Test-TimeCompute with Latent…Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth ApproachdLLM-Cache: AcceleratingDiffusion Large Languag…dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive CachingDiffusion LanguageModels Know the Answer…Diffusion Language Models Know the Answer Before DecodingLLaDA-MoE: A Sparse MoEDiffusion Language ModelLLaDA-MoE: A Sparse MoE Diffusion Language ModeldParallel: LearnableParallel Decoding for…dParallel: Learnable Parallel Decoding for dLLMsEfficient-DLM: FromAutoregressive to…Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in SpeedCoevolutionaryContinuous Discrete…Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent ReasonerSoft-Masked DiffusionLanguage ModelsSoft-Masked Diffusion Language ModelsContinuously AugmentedDiscrete Diffusion mode…Continuously Augmented Discrete Diffusion model for Categorical Generative ModelingCANDI: HybridDiscrete-Continuous…CANDI: Hybrid Discrete-Continuous Diffusion ModelsFast Solvers forDiscrete Diffusion…Fast Solvers for Discrete Diffusion Models: Theory and Applications of High-Order AlgorithmsdLLM: Simple DiffusionLanguage ModelingdLLM: Simple Diffusion Language ModelingPrism: EfficientTest-Time Scaling via…Prism: Efficient Test-Time Scaling via Hierarchical Search and Self-Verification for Discrete Diffusion Language ModelsDream 7B: DiffusionLarge Language ModelsDream 7B: Diffusion Large Language Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。