Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding

Diffusion-based large language models (Diffusion LLMs) have shown promise for non-autoregressive text generation with parallel decoding capabilities. However, the practical inference speed of open-sourced Diffusion LLMs often lags behind autoregressive models due to the lack of Key-Value (KV) Cache and quality degradation when decoding multiple tokens simultaneously. To bridge this gap, we introduce a novel block-wise approximate KV Cache mechanism tailored for bidirectional diffusion models, enabling cache reuse with negligible performance drop. Additionally, we identify the root cause of generation quality degradation in parallel decoding as the disruption of token dependencies under the conditional independence assumption. To address this, we propose a confidence-aware parallel decoding strategy that selectively decodes tokens exceeding a confidence threshold, mitigating dependency violations and maintaining generation quality. Experimental results on LLaDA and Dream models across multiple LLM benchmarks demonstrate up to \textbf{27.6$\times$ throughput} improvement with minimal accuracy loss, closing the performance gap with autoregressive models and paving the way for practical deployment of Diffusion LLMs.

Attention Is All YouNeedAttention Is All You NeedDiffusionBERT: ImprovingGenerative Masked…DiffusionBERT: Improving Generative Masked Language Models with Diffusion ModelsScore-basedContinuous-time Discret…Score-based Continuous-time Discrete Diffusion ModelsDiffusion LanguageModels Can Perform Many…Diffusion Language Models Can Perform Many Tasks with Scaling and Instruction-FinetuningDiscrete DiffusionLanguage Modeling by…Discrete Diffusion Language Modeling by Estimating the Ratios of the Data DistributionSimplified andGeneralized Masked…Simplified and Generalized Masked Diffusion for Discrete DataSimple and EffectiveMasked Diffusion…Simple and Effective Masked Diffusion Language ModelsFast Sampling viaDiscrete Non-Markov…Fast Sampling via Discrete Non-Markov Diffusion Models with Predetermined Transition TimeYour Absorbing DiscreteDiffusion Secretly…Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean DataMasked Diffusion Modelsare Secretly…Masked Diffusion Models are Secretly Time-Agnostic Masked Models and Exploit Inaccurate Categorical SamplingScaling DiffusionLanguage Models via…Scaling Diffusion Language Models via Adaptation from Autoregressive ModelsLLaDA-V: Large LanguageDiffusion Models with…LLaDA-V: Large Language Diffusion Models with Visual Instruction TuningFast-dLLM v2: EfficientBlock-Diffusion LLMFast-dLLM v2: Efficient Block-Diffusion LLMAccelerating DiffusionLarge Language Models…Accelerating Diffusion Large Language Models with SlowFast Sampling: The Three Golden PrinciplesDiffusion LanguageModels Know the Answer…Diffusion Language Models Know the Answer Before DecodingRevolutionizingReinforcement Learning…Revolutionizing Reinforcement Learning Framework for Diffusion Large Language ModelsdParallel: LearnableParallel Decoding for…dParallel: Learnable Parallel Decoding for dLLMsLLaDA-MoE: A Sparse MoEDiffusion Language ModelLLaDA-MoE: A Sparse MoE Diffusion Language ModelLearning UnmaskingPolicies for Diffusion…Learning Unmasking Policies for Diffusion Language ModelsdUltra: Ultra-FastDiffusion Language…dUltra: Ultra-Fast Diffusion Language Models via Reinforcement LearningAdaBlock-dLLM:Semantic-Aware Diffusio…AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block SizeDMax: AggressiveParallel Decoding for…DMax: Aggressive Parallel Decoding for dLLMsd3LLM: Ultra-FastDiffusion LLM using…d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory DistillationLongLLaDA: UnlockingLong Context…LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMsFast-dLLM: Training-freeAcceleration of…Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。