Trust but Verify: Designing Region-Inhomogeneous Forward Dynamics for Discrete Diffusion
Abstract
In continuous diffusion, the forward process is a design surface: noise schedules, SDE variants, and per-position noise levels are all engineered to match the task. Discrete diffusion for language has collapsed this surface to a single point, the ab- sorbing process, under which conditioning context sits at infinite signal-to-noise ratio for the entire reverse trajectory. The collapse has a cost. Trust in the con- ditioning becomes binary, so corrupted context tokens are copied as ground truth and their errors propagate into the generated spans. We treat the discrete forward process as a design space again, composing absorbing and uniform-mixing cor- ruption kernels with region-dependent rates: hole positions traverse the standard absorbing dynamics, while user-supplied context is corrupted by light random sub- stitution at a small rate η, so that “observed but uncertain” is a state the model has visited during training. At inference, a heuristic reverse-time sampler re-injects noise into suspicious context tokens under an annealed intensity (a corrector-like step that acts on the conditioning rather than the sample) and freezes the context dynamics late in the trajectory. An 8,000-step finetune of a 170M-parameter pre- trained masked diffusion model, under five hours on a 6 GB laptop GPU, converts a blind copier into a verifier. The base model assigns roughly 50% probability to observed tokens even when they are uniform-random garbage; on conspicu- ous (uniform-random) corruptions, the finetuned model detects corrupted context with 0.96 recall at 0.91 precision, improves infill token-F1 by 3.8 points at 10% context corruption (6.4 at 20%), and eliminates the accuracy dip in tokens gener- ated adjacent to corruptions, with no measurable regression on clean conditioning (false-edit rate under 0.4% on clean context). Gains under plausible corruptions are much weaker (about 0.5 F1 points at 10% corruption, not statistically signif- icant). The experiments also suggest a regularity: detectability and damage are jointly governed by how far the corruption kernel departs from the data condi- tional. Noise near the conditional is both statistically invisible and semantically light; noise far from it is conspicuous and destructive. We interpret absorbing masked diffusion as the infinite-contrast limit of a trust axis whose continuous analog is the SNR gap between known and unknown regions in diffusion inpaint- ing; in our setting (one 170M model, one text domain), points inside that axis can be reached by perturbing the forward dynamics rather than the architecture.