Loopholing All the Way: Continuous Looping with Discrete Diffusion Models
Abstract
Discrete Diffusion Models (DDMs) offer an alternative to autoregressive models, and support parallel text generation. However, when DDM samplers update multiple positions in a single step, they traditionally sample tokens independently, conditioned on the current noisy sequence. This limits their ability to capture dependencies among tokens sampled together and hurt the generation quality when using few steps. Yet LDDMs still sample discrete tokens at each denoising step, so they still sample tokens independently from categorical distributions. Recent fully continuous Flow Language Models (FLMs) avoid factorized sampling, but lag behind DDMs, especially on math and code. We propose \method (``Loopholing All the Way''), which converts a pretrained DDM into a continuous generative model with cheap fine-tuning. LAW feeds soft embeddings back into the denoising backbone and samples from the categorical distribution only at the final step, as in recent FLMs. With ~20k fine-tuning steps on a single GPU, LAW increases the GSM8K accuracy of Duo, a strong uniform DDM, from ~18% to ~55%, approaching the performance of ARMs. In few-step generation, LAW reaches ~46\% accuracy in 16 steps, whereas strong distillation methods need 128 steps to reach ~21%.