\texttt{FEROM}: Frontier Endogenous Reveal-Order Marginal Policy Optimization for Masked Diffusion LMs
Abstract
Masked diffusion language models generate text by iteratively unmasking positions, with practical samplers often selecting reveal positions from the model's own logits. The reveal order is therefore an endogenous latent variable of the rollout process with an enormous discrete latent space. Policy optimization objectives relying on pre-defined schedules or sampled trajectories introduce either bias or high variance. In this paper, we propose \texttt{FEROM}, \emph{Frontier Endogenous Reveal-Order Marginal Policy Optimization}, which targets the rollout-induced marginal response policy of masked diffusion LMs. \texttt{FEROM} derives a Rao--Blackwellized policy-gradient identity over latent reveal orders and expresses the resulting estimator as posterior edge occupancy on a reveal-state DAG. To make marginalization practical, we introduce Frontier Reveal Marginalization, a budgeted estimator that combines scorelaw local reveal estimation with frontier expansion of high-mass partial states. Integrated into a GRPO-style objective, \texttt{FEROM} replaces single-path log-scores with a locally marginalized edge-based surrogate. Experiments on math and coding tasks show comparable or improved results over existing methods under matched compute budgets. Offline proxy study further shows potential in gains with increased budgets.