Diffusion Language Models Are Natively Length-Aware
Vittorio Rossi ⋅ Giacomo Cirò ⋅ Davide Beltrame ⋅ Luca Gandolfi ⋅ Paul Röttger ⋅ Dirk Hovy
Abstract
Diffusion Language Models (DLMs) operate over a fixed-length canvas for a predetermined number of denoising steps. This process is independent of the actually required response length, so short responses (typical of chat tasks) waste compute on denoising empty end-of-sequence (EOS) tokens: the padding tax. Recent work has tried to address this issue in various ways, but all treat the EOS as a flawed content token the sampler must compensate for. Instead, we show that what prior work identified as a flaw to be patched is a well-defined signal that was being misused. DLMs trained with EOS-as-padding are already natively length-aware: under the standard masked-diffusion objective, the optimal EOS probability at every position and noise level coincides with the conditional CDF of the response length. That allows us to recover the entire prompt-conditional length CDF. In other words: an optimal model knows the response length at the first diffusion step. We implement this principle across three DLMs via SmartCrop, a minimal training-free decoder that reads the CDF once, crops the canvas at a target quantile, and then denoises the cropped window. Across four math and code benchmarks, SmartCrop matches recent approaches on 15 of 17 paired tests despite reading the CDF only once, evidence that the length signal is available at $t=1$.
Chat is not available.
Successful Page Load