Empowering Masked Diffusion Models to Self-Correct with Leave-One-Out Transformers
Abstract
Diffusion models promise iterative refinement, yet masked diffusion language models (MDMs) suffer from two shortcomings: (1) inability to revise tokens, and (2) sparse training signal. The root cause is structural: MDMs control information flow only through input masking, so unmasked positions cannot be supervised, and self-correction is never learned. We propose the Leave-One-Out Transformer (LOOT), which addresses both issues. With LOOT, MDMs predict the distribution over potential replacements conditioned on the rest of the sequence at each position, in a single forward pass. Supervision can then be applied at masked and unmasked positions alike, yielding a lower-variance training objective with a dense loss signal at every token. Our method empowers MDMs to learn self-correction during training and supports continuous refinement at inference time via Gibbs sampling. LOOT achieves a new best validation perplexity on LM1B in a parameter-matched setting. Finetuned from a MDM checkpoint on OpenWebText, LOOT establishes a generative quality-vs-diversity frontier that surpasses prior remasking methods across NFE budgets.