How Much Dependency Can a Marginal Rule See? Total Correlation in Diffusion LM Parallel Decoding
Dovud Asadov ⋅ Rustambek Urokov ⋅ Qobiljon Toshnazarov ⋅ KAYUMOV ABDUAZIZ ⋅ Mardon Hazratov ⋅ Husanboy Mansuraliyev
Abstract
A masked diffusion language model commits several masked positions per step. When its conditionals compose into a joint, the exact cost of committing a set $S$ is that joint's conditional total correlation $C(S\mid c)$. When they do not, we measure the order-symmetric estimate and report the discrepancy alongside it. Many systems choose the commit set from per-position marginals. That such a rule cannot determine $C$ is elementary, and we make it exact: the summed marginal entropy $B$ certifies $C\le B$, but at fixed marginals $C$ realises every value in $[0,C_{\max}]$ with $C_{\max}\le(1-1/k)B$, exactly $k-1$ of $k$ bits for fair bits. The worst case is not the practitioner's question, though. We ask instead how much of $C$ is recoverable from the marginals on a real checkpoint, and answer it by estimating $C$ per candidate set on four masked diffusion LMs across two families, with $247$ to $516$ blocks per model and regime. In practice, the marginals locate $C$ far more tightly than the worst case permits: where the bound allows a window $2.6$ to $13.0$ bits wide, a marginal-only predictor scored out-of-fold and corrected for estimator variance leaves a residual of $0.74$ to $1.17$ bits on six of eight model-regime settings ($1.62$ and $1.84$ on the one checkpoint whose conditionals are least self-consistent). Getting there requires a readout one can trust, and that is not a formality: Dream and LLaDA need opposite logit-position conventions, and reading Dream with the wrong one costs $6.9$ bits on the correct token, inflates a conditional-inconsistency diagnostic by $14$ to $41\times$, and flips the sign of the confidence-error correlation on both Dream checkpoints, a clean and entirely spurious finding that an earlier version of this work reported. We therefore specify three gates that measure the convention rather than assume it, and report their output for every checkpoint. Finally, scoring all six marginal rules against measured $C$ contradicts our own synthetic audit: pooling across the block beats gating on its weakest position in every setting, and the top-2 margin, which dominates confidence on tasks with analytic $C$, loses to it on real models.
Chat is not available.
Successful Page Load