Free Draft-and-Verification: Toward Lossless Parallel Decoding for Diffusion Large Language Models
Shutong Wu ⋅ Jiawei zhang
Abstract
Diffusion Large Language Models (dLLMs) have emerged as a compelling paradigm for language modeling, offering the unique potential for highly efficient inference via multi-token parallel decoding. However, unlocking this potential remains challenging. High-quality generation typically defaults to one-token-per-step static decoding, while existing parallel algorithms rely on heuristic confidence thresholds, introducing a fragile trade-off between decoding efficiency and generation quality. To overcome this, we introduce **Free** **D**raft-**a**nd-**Ve**rification (**FreeDave**), a novel training-free and model-free fast decoding algorithm. FreeDave capitalizes on the inherent properties of dLLMs by treating future token predictions as "free drafts" and rigorously self-verifying them in subsequent forward passes. Theoretically, FreeDave is guaranteed to use the fewest possible model forward passes to reproduce the exact sequence generated by static decoding, achieving algorithmically lossless acceleration. Extensive evaluations on math reasoning and code generation benchmarks demonstrate that FreeDave can boost the number of generated tokens per model forward (TPF) by up to 5.48$\times$ and tokens per second (TPS) by up to 3.94$\times$, while robustly maintaining generation quality and avoiding the task-specific performance degradation faced by heuristic methods.
Chat is not available.
Successful Page Load