D-DOIT: Training-free Adaptation of Discrete Diffusion via Doob's h-Transform
Jieke Wu ⋅ Qijie Zhu ⋅ Weimin Wu ⋅ Zeqi Ye ⋅ Minshuo Chen ⋅ Han Liu
Abstract
We propose D-DOIT (Discrete Doob-Oriented Inference-time Transformation), a training-free and efficient adaptation method for discrete diffusion models with generic rewards. D-DOIT formulates adaptation as sampling from a reward-tilted target distribution and realizes this transport through Doob's h-transform of the discrete diffusion reverse kernel, using only reward values rather than reward gradients. Unlike continuous diffusion, masked discrete diffusion samples categorical token-reveal transitions rather than continuous state updates. D-DOIT derives the corresponding discrete Doob's h-transform, which guides sampling by reweighting reverse transition probabilities instead of adding a drift correction. To make this transformation practical, D-DOIT avoids expensive future rollouts. At each guided step, D-DOIT samples candidate next states, uses the model prediction head to complete each candidate into a clean sequence, evaluates each completion with the reward oracle, and resamples the next state with probabilities proportional to the rewards. An optional late-stage best-of-$K$ refinement further improves sample quality by branching trajectories only near the end of denoising, avoiding the $K$-fold cost over the full trajectory. Empirically, on regulatory DNA sequence design benchmarks, D-DOIT consistently outperforms training-free guidance baselines and achieves performance competitive with training-based methods, improving both enhancer activity and cell-type-specificity while preserving sequence naturalness.
Chat is not available.
Successful Page Load