How Reliable are Intermediate Rewards for SMC-based Guidance in Masked Discrete Diffusion Models? A Closed-Form Analysis
Abstract
Sequential Monte-Carlo (SMC) methods offer an appealing approach for guided generation in masked discrete diffusion models as they do not require gradient information, and can be applied to pre-trained models without retraining. SMC guides generation by assigning rewards to partially generated samples, and we investigate the impact of common choices for computing such rewards. We formulate a biologically relevant optimisation problem as a masked discrete diffusion process, yielding all required quantities exactly, cheaply, and in closed-form, eliminating approximation error. Using this setup, we identify an issue with serious implications: for rewards which yield no information for infeasible sequences, as is often the case in scientific problems, scoring partially generated samples based on samples from the factorised distribution, as is commonly done, results in poor performance. Conversely, these cheap samples can lead to effective guidance if the reward yields information about infeasible sequences. Our two conclusions suggest that we may need to rethink how to allocate inference-time compute for successful guidance in challenging scientific problems.