Right for the Wrong Reasons: Interval Quality, Not Directional Accuracy, Separates Foundation Models on Intraday Returns
Abstract
Diffusion language models decode by iteratively unmasking tokens, raising the possibility of stopping once the answer has stabilized rather than paying for a fixed denoising budget. We study when stable-correct convergence occurs, whether it is detectable online, and whether detection produces a wall-clock benefit. On all 1,319 GSM8K test problems with 128 deterministic denoising steps, we define the hindsight-optimal convergence step (HOCS) as the earliest step whose all-in completion is correct at that step and every later step. LLaDA-8B-Instruct reaches HOCS on 54.74% of problems at a conditional mean step of 47.98, yielding a 1.520x zero-additional-error HOCS-oracle step speedup; Dream-7B reaches HOCS on 38.06% at a conditional mean step of 119.84, yielding only 1.025x. Linear probes over pooled hidden states and progress separate safe from unsafe steps in both models (AUROC 0.879 and 0.987); removing the explicit progress coordinate leaves ranking performance unchanged to three decimals under the fixed schedule. Fresh online inference confirms that detectability and exploitable headroom are distinct. On LLaDA, probe plus stability retains 97.78% of full-decoding accuracy while reaching 1.588x step and 1.384x latency speedup. On Dream, the same policy retains 100.20% but reaches only 1.016x step and 0.996x latency speedup. Retention can exceed 100% because an early commit occasionally corrects an answer that full decoding later changes. On MATH-500, frozen-probe accuracy differs from full decoding by at most 0.2 percentage points, but no probe policy provides a latency benefit. Detector quality, convergence timing, stopping risk, and signal cost must therefore be evaluated jointly: a detectable state is not necessarily an actionable one.