Iterative Value-Guided Tilting for Fine-Tuning Masked Diffusion Models: An Actor–Critic Approach
Abstract
Discrete diffusion models have demonstrated considerable success in various generative tasks, ranging from natural language to biological sequences. However, fine-tuning these models on fine-grained black-box rewards such as binding affinity or molecular property scores remains a challenging task. We propose a novel algorithm for masked diffusion models that requires only black-box reward evaluations and follows an actor-critic-like procedure, alternating between learning a value function from current-model samples and optimizing the generative model using the resulting value estimates. We validate our method on DNA sequence design and protein inverse folding tasks, outperforming comparable methods that rely only on black-box reward evaluations, while remaining competitive with approaches that additionally exploit reward gradients or labeled data. Using a simple two-dimensional example, we also illustrate the benefit of iterating between value estimation and model optimization, rather than performing a single optimization round. Our results suggest that actor-critic-style optimization is a practical approach to fine-tuning masked diffusion models under black-box rewards.