Why Deterministic PRM Guidance Underperforms in Discrete Diffusion Reasoning
Yan Zhan ⋅ Shaobo Liu ⋅ Zhijun Gao
Abstract
We present a diagnostic study of reward guided reasoning in discrete diffusion language models (dLLMs). On in-distribution GSM8K with Dream-v0-Instruct-7B, deterministic segmental top-$1$ guidance using outcome-supervised intermediate-state reward models underperforms a simpler matched-compute baseline: independent best-of-$N$ sampling plus a GSM8K-trained final-answer outcome reward model (ORM) reranker. The setup appears well suited for PRM guidance because dLLM intermediate states partially expose solution evidence; however, these states are unordered masked subsets rather than autoregressive prefixes, and prior reward-guided decoding results rarely report matched forward-pass compute against simple verifier reranking. We study \emph{outcome-supervised} PRMs trained with binary final-correctness labels on dLLM intermediate states. Under a protocol that charges denoising, PRM scoring, and ORM scoring in the same forward-pass unit, deterministic PRM-guided search trails ORM Rerank by $9.95$ and $12.69$ percentage points (pp) in seed-mean aggregate at matched candidate budgets $K{=}8$ and $K{=}32$; paired bootstrap CIs on seed $42$ exclude zero. Two mechanisms account for much of the headline gap: bidirectional PRM ROC-AUC (area under the receiver operating characteristic curve) decays from $0.77$ to $0.54$ with mask ratio, and deterministic top-$1$ pruning drops the guided pool's perfect-selector ceiling by $\sim$$14$\,pp. A third diagnostic shows that mean pooling also weakens causal PRM variants. Matched-compute ORM Rerank thus emerges as the strong baseline when an in-distribution final-answer verifier is available; closing the gap requires PRM guidance that preserves candidate diversity and queries the scorer at denoising stages where it remains discriminative, neither of which the deterministic top-$1$ recipe satisfies.
Chat is not available.
Successful Page Load