Renoise Consistency: Unlocking Efficient Self-Correction for Diffusion Large Language Models
Abstract
Diffusion Large Language Models (DLMs) are being actively explored as promising alternatives to autoregressive models due to their fast and flexible generation capabilities. To optimize inference in DLMs, recent studies have proposed a variety of decoding algorithms. However, the tendency of these methods to mimic autoregressive generation limits the speed and flexibility of DLMs, while also preventing the methods from fully realizing their potential. Moreover, this rigid approach leads to error propagation, while existing sampling-based correction methods incur prohibitive computational costs. To overcome this, we propose Renoise Consistency, a novel post-hoc approach that enables effective self-correction by maximizing train--inference consistency and can be applied to any decoding algorithm with a flexible computational budget. Furthermore, we introduce Adaptive Block Drafting, which leverages consistent internal representations of DLMs to significantly reduce the overall computational cost. When combined, our proposed method achieves up to a 9.00\% performance improvement over the semi-autoregressive baseline, along with a 25.85\% reduction in forward steps.