GoldiMask: Fine-Tuning Diffusion Language Models in the Measured Zone of Learnability
Abstract
Masked diffusion language models perform supervised fine-tuning by randomly masking tokens and learning to reconstruct them from the remaining context. However, not every mask choice is equally useful for training. A masked token whose value directly depends on another masked one can only be memorized, and one already obvious from context does not yield a useful gradient. Selecting a good mask, therefore, means answering two questions: (i) which tokens are worth supervising now, and (ii) which of the others would most help to be used for supervision. To address them, we propose two diagnostic curves that are efficiently computable with the base model. First, the \emph{learnability curve} shows that the benefit of supervising a token peaks at intermediate gold token probability and bottoms out at both extremes. Second, the \emph{support curve}, shows that the revealed context helps in proportion to the attention a masked token directs at this revealed context, with diminishing returns. Based on these diagnostics, we introduce \textbf{GoldiMask}, a training strategy that replaces the random masking split with the maximizer of a single objective built from the two curves: at the sampled masking ratio, it scores a split by how much support the revealed tokens give to the positions left masked, weighted by how much those positions can still learn. The objective is submodular, and self-paying by construction --- revealing a token forfeits that token's own supervision term, so every reveal is priced against what it costs. GoldiMask then grades each supervised token by the learning it actually \emph{realized}, measured as the increase in gold-token probability from the diagnostic pass to the supervised pass, and divides each training step's fixed loss-weight budget among tokens by a provably proportionally-fair rule. Against baselines retrained in a single harness, GoldiMask achieves state-of-the-art results across three models and two datasets, especially on benchmarks that require a chain of reasoning to be reconstructed rather than memorized.