Discrete Stochastic Localization for Non-autoregressive Generation
Yunshu Wu ⋅ Jiayi Cheng ⋅ Longxuan YU ⋅ Partha Thakuria ⋅ Rob Brekelmans ⋅ Vagelis Papalexakis ⋅ Greg Ver Steeg
Abstract
Continuous diffusion is a natural framework for non-autoregressive generation but has consistently underperformed masked discrete diffusion (MDM) on discrete sequence generation. We argue this is not a limitation of continuity itself, but of how previous models parameterize the denoiser as a family of timestep-indexed regimes. We introduce \emph{Discrete Stochastic Localization} (DSL), a continuous-state framework with unit-sphere token embeddings whose Bayes-optimal denoiser is invariant to the nominal signal-to-noise ratio (SNR). One trained network then supports an entire family of per-token SNR paths, with masked diffusion as one special case. Fine-tuning a pretrained MDLM checkpoint with DSL substantially improves distributional faithfulness (MAUVE) on OpenWebText across all step budgets from $T{=}128$ to $T{=}1024$, and the same checkpoint supports random-order autoregressive sampling and a hybrid continuous-then-discrete sampler at as few as $T{=}48$ total steps---without distillation or retraining. On Text8, DSL gives the first continuous-state NLL estimate approaching the MDM range. Code is anonymously available at \url{https://anonymous.4open.science/r/DSL_anonymous-94B5/}.
Chat is not available.
Successful Page Load