ConGO: Consistency-guided Optimization for Parallel Decoding in Discrete Diffusion Models
Abstract
Discrete diffusion is an emerging paradigm for generative AI, in which a sequence of tokens is generated through iterative updates using a learned denoising model. While the standard formulation randomly selects token positions to update in each denoising iteration, common heuristics select the positions with highest marginal confidence or minimum inter-token dependency to increase generation quality, particularly in parallel decoding. These approaches, however, strictly isolate the position selection from the content selection. Such value-blind position selection hinders the commitment of jointly consistent tokens, which can be indicated by explicit constraints or learned priors. In this work, we propose consistency-guided optimization (ConGO) for parallel decoding in discrete diffusion models, which formulates the update selection as a quadratic unconstrained binary optimization problem. Guided by a lightweight, value-aware consistency model, ConGO offers a flexible spectrum of integration depending on the reliability of the available knowledge. When strong, explicit domain knowledge is available, such as in Sudoku, ConGO jointly selects both the positions and their content, achieving significantly higher accuracy on hard problems (up to a 44 pp improvement). When explicit constraints are absent, ConGO restricts its intervention purely to position selection based on learned pairwise priors (e.g., bigrams), which substantially reduces LLaDA-8B-Instruct's required number of function evaluations on math and coding (by 25% on GSM8K and 42% on HumanEval) while maintaining iso-accuracy.