Refinement Buys Intelligibility, Search Buys Identity: What Test-Time Compute Buys in Masked-Diffusion TTS
Abstract
A diffusion language model buys sequential computation twice: depth, paid once in parameters, and denoising steps, rented per utterance at serve time. We ask whether the rented kind helps every capability equally. We train an iso-parameter width–depth grid of 15 masked-diffusion codec TTS models (3 seeds, 19–133M non-embedding parameters) on 2,000 h of Emilia-EN and sweep the refinement budget (T) at inference only, evaluating cross-sentence zero-shot synthesis on 400 items from 174 held-out speakers. Our central result is model-free and referenced to floors we measure rather than to zero: the word error of an ASR system on the real recordings, and the speaker similarity of a codec round trip. Against these floors, refinement from (T=1) to (T=16) closes 86.2% of the reachable intelligibility range but only 46.4% of the reachable identity range, a 1.86× asymmetry that holds between 1.68× and 1.86× under every monotone reparameterisation of the error we tried. Retraining for 3× as long narrows the factor to 1.36× while preserving its direction, so we offer the ordering, not the factor, as the result. We also report the pre-registered power-law fit, and we report that the exponent contrast it yields ((\Delta\tau = \tau_{\mathrm{WER}} - \tau_{\mathrm{SIM}} = 0.110)) changes sign under those same reparameterisations: it is a property of the error coordinate, and we therefore decline to build the paper on it. The limit is on what refinement buys, not on what inference buys. At matched NFE, best-of-(K) search moves speaker similarity where refinement cannot, and four speaker-encoder families outside the selector’s lineage agree, with per-item win rates of 64.6–79.0% and every interval above one half. Depth and steps are not interchangeable: the substitution form (B(dT^\kappa)^{-\beta}) fits worse than the separable form by (\Delta\mathrm{AICc}=+69.3). Finally, 62% of the remaining identity gap belongs to the codec, not the model—so most of what identity has left to gain is not the generator’s to win.