Return-Weighted ELBO Fine-Tuning Degrades Masked Diffusion Planners
Abstract
Fine-tuning a masked discrete diffusion planner with a return-weighted ELBO makes it worse. We run 25 conditions with three seeds each on Craftax Classic and MiniHack. The claim rests on Craftax Classic, where no condition recovers the checkpoint it started from. There, collected return falls alongside held-out score, suggesting that this is not reward hacking. The conditions that lose least there are LoRA, a hard KL trust region and head-only updates, which barely train. We test whether the return signal is responsible, using an exact decomposition of the gradient into an imitation term and a return term, both directly measurable. On Craftax Classic the return term is large, about half the imitation gradient in norm. But shrinking it does not help. Advantage clipping cuts the return term fivefold and scores \emph{below} the unclipped baseline. Unweighted ELBO on all self-generated rollouts degrades further still on both benchmarks, so the weighting looks protective. What survives is an ordering by how far each condition lets the weights move. Score rises \emph{with} the size of the return term, and that ordering tracks the weights' dispersion rather than the return. Continued imitation also degrades the checkpoint, by less than half as much, and collecting at the evaluation settings removes most of the degradation for both. Most of it therefore lies in the sampling mismatch rather than the return weighting, on one benchmark and one model. We give both diagnostics; each costs at most one batch and no accelerator.