When Do Denoising Errors Propagate? A Theory of Downstream Sensitivity in Diffusion Language Models
Preeti ⋅ Venkata S Dhara ⋅ KIRAN RAVISH ⋅ Ankita Kushwaha ⋅ Pawan Kumar
Abstract
Diffusion language models generate through iterative reverse dynamics, so a local denoising error may be revised or remasked while still influencing later predictions. We formalize this phenomenon through \emph{continuation sensitivity}, separating local error likelihood from downstream consequence. For general finite-state reverse kernels, we derive a Hamming-Wasserstein stability bound decomposing final discrepancy into local approximation error, inference-to-reference occupancy shift, and future sensitivity. For parallel decoding, we show that high confidence and negligible within-batch dependence do not prevent an $O(\delta)$ perturbation from causing $\Theta(\delta n)$ expected Hamming damage. Under fixed irreversible schedules, a coordinate-influence certificate motivates a practical estimator $\widehat{\Lambda}$. Across $45{,}600$ controlled interventions on OpenWebText, GSM8K, and MBPP with MDLM, Qwen3-0.6B-diffusion, and LLaDA-8B-Instruct, higher estimated influence is associated with greater final damage at matched confidence and reverse time in all five settings ($d=0.175$-$0.781$). The association persists after controlling for an empirical within-batch dependency proxy, although incremental predictive value is strongly model-dependent. These results establish downstream sensitivity as a complementary axis of DLM decoding risk.
Chat is not available.
Successful Page Load