What Does a Discrete Diffusion Model Learn? One Reverse Process in Three Coordinates
Abstract
The map from a discrete-diffusion model's output to the denoising process is described in the literature in three ways: as a denoiser, a bridge plug-in, or a concrete score. We show that these are three coordinates of the same reverse process and derive the exact optimizer of each of their training objectives. We do so via the Oracle Distance identity: the negative continuous-time ELBO is not just a bound on the likelihood but the data entropy plus the path KL from the optimal process to the learned one, obtaining its unique, general optimizer and showing that all processes share the same best attainable negative ELBO. Computing this optimizer for token-factorizing noise leads to the three coordinates used in the literature, with closed-form conversions, and recovers all the standard formulas. In particular, the bridge plug-in does not train the denoiser law but the cavity law: the law of a clean token given its context but excluding its own noisy token; for masked diffusion the two coincide, which is why the distinction had remained unnoticed. We validate these results and demonstrate the relevance of the cavity-to-denoiser distinction at language-modeling scale, comparing for the first time all three coordinates across masked, uniform, and GIDD diffusion under a shared recipe. With further results, such as ELBO calibration identities and the equivalence of time schedules and training weights, this gives a unified framework and a practical recipe for designing and comparing new kernels.