Large Language Diffusion Models

The capabilities of large language models (LLMs) are widely regarded as relying on autoregressive models (ARMs). We challenge this notion by introducing LLaDA, a diffusion model trained from scratch under the pre-training and supervised fine-tuning (SFT) paradigm. LLaDA employs a forward data masking process and a reverse generation process, parameterized by a Transformer to predict masked tokens. It provides a principled generative approach for probabilistic inference by optimizing a likelihood lower bound. Across extensive benchmarks on general tasks, math, code, and so on, LLaDA demonstrates strong scalability and performs comparably to our self-constructed ARM baselines. Remarkably, LLaDA 8B is competitive with strong LLMs like LLaMA3 8B in in-context learning and, after SFT, exhibits impressive instruction-following abilities in case studies such as multi-turn dialogue. Moreover, LLaDA addresses the reversal curse, surpassing GPT-4o in a reversal poem completion task. Our findings show the promise of diffusion models for language modeling at scale and challenge the common assumption that core LLM capabilities discussed above inherently depend on ARMs. Project page and codes: https://ml-gsai.github.io/LLaDA-demo/.

Continuous diffusion forcategorical dataContinuous diffusion for categorical dataDiffusion LanguageModels Can Perform Many…Diffusion Language Models Can Perform Many Tasks with Scaling and Instruction-FinetuningDiscrete DiffusionLanguage Modeling by…Discrete Diffusion Language Modeling by Estimating the Ratios of the Data DistributionSimple and EffectiveMasked Diffusion…Simple and Effective Masked Diffusion Language ModelsSimplified andGeneralized Masked…Simplified and Generalized Masked Diffusion for Discrete DataDeepSeek LLM: ScalingOpen-Source Language…DeepSeek LLM: Scaling Open-Source Language Models with LongtermismScaling up MaskedDiffusion Models on TextScaling up Masked Diffusion Models on TextBlock Diffusion:Interpolating Between…Block Diffusion: Interpolating Between Autoregressive and Diffusion Language ModelsYour Absorbing DiscreteDiffusion Secretly…Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean DataScaling DiffusionLanguage Models via…Scaling Diffusion Language Models via Adaptation from Autoregressive ModelsMercury: Ultra-FastLanguage Models Based o…Mercury: Ultra-Fast Language Models Based on DiffusionMasked Diffusion Modelsare Secretly…Masked Diffusion Models are Secretly Time-Agnostic Masked Models and Exploit Inaccurate Categorical Samplingd1: Scaling Reasoning inDiffusion Large Languag…d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement LearningLLaDA-MoE: A Sparse MoEDiffusion Language ModelLLaDA-MoE: A Sparse MoE Diffusion Language ModelFast Solvers forDiscrete Diffusion…Fast Solvers for Discrete Diffusion Models: Theory and Applications of High-Order AlgorithmsAccelerating DiffusionLarge Language Models…Accelerating Diffusion Large Language Models with SlowFast Sampling: The Three Golden PrinciplesEfficient-DLM: FromAutoregressive to…Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in SpeedFine-Tuning MaskedDiffusion for Provable…Fine-Tuning Masked Diffusion for Provable Self-CorrectiondParallel: LearnableParallel Decoding for…dParallel: Learnable Parallel Decoding for dLLMsDiffusion BeatsAutoregressive in…Diffusion Beats Autoregressive in Data-Constrained SettingsTest-Time Alignment ofDiscrete Diffusion…Test-Time Alignment of Discrete Diffusion Models with Sequential Monte CarloA Convergence Theory forDiffusion Language…A Convergence Theory for Diffusion Language Models: An Information-Theoretic PerspectiveEncoder-DecoderDiffusion Language…Encoder-Decoder Diffusion Language Models for Efficient Training and InferencePrism: EfficientTest-Time Scaling via…Prism: Efficient Test-Time Scaling via Hierarchical Search and Self-Verification for Discrete Diffusion Language ModelsLarge Language DiffusionModelsLarge Language Diffusion ModelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.