When to Stop and Which to Return: Trajectory-Level Minimum Bayes Risk Decoding for Diffusion Language Models
Subeen Park ⋅ Sungjun Lim ⋅ Heejin Jung ⋅ Somin Kim ⋅ Gyeong-Moon Park ⋅ Kyungwoo Song
Abstract
Masked diffusion language models can produce high-quality intermediate responses before denoising finishes, yet continued decoding may incur unnecessary computation or overwrite better responses. Existing approaches address these challenges separately. Early-commitment methods reduce computation but cannot recover earlier responses. Temporal voting exploits intermediate responses to improve quality but retains the cost of full decoding. We propose **Trajectory-MBR**, a training-free method that jointly decides when to stop and which current or earlier response to return. It reuses denoiser outputs to construct complete response candidates and performs trajectory-score-weighted minimum Bayes risk selection within a sliding window. Once the minimum weighted risk falls a threshold, decoding stops and returns the selected response. We show that accurate trajectory-based utility estimates yield near-optimal selection within the window and provide an offline condition for improvement over full decoding. Across language and multimodal tasks, Trajectory-MBR maintains or improves average response quality with average denoiser-evaluation speedups of up to $3.76\times$, while its token-level extension preserves quality at a $12.34\times$ speedup on HumanEval.
Chat is not available.
Successful Page Load