Separating Rate from Signal in Per-Token Masking Schedules in Diffusion Language Models
Övgü Özdemir ⋅ Erdem Akagündüz
Abstract
Masked diffusion language models corrupt text by replacing tokens with a mask symbol, and in the standard formulation every token is masked with the same probability. Several methods make that probability depend on the token through a per-token signal, such as corpus frequency. Such a schedule changes three things at once: the average masking rate, its spread across tokens, and which token receives which weight. Only the last depends on the signal, and reported gains are not attributed among them. We give three matched controls that isolate each property, a parameterization under which the mean cannot drift, and a check, before training, that the weights order tokens as intended. Applying these to an instance built from frozen BERT similarity as a per-token signal, the weights do not rank importance, placing a content word above a function word in $30.7\%$ of pairs where chance is $50\%$, and their mean of $1.178$ makes the schedule mask fewer tokens than the baseline. On WikiText-103 the instance improves perplexity by $2.8\%$; a constant weight carrying no token-level information reproduces $41.2\%$ of that, permuting the weights reproduces $30.2\%$, and removing the mean shift while keeping the ordering leaves perplexity within seed noise of the baseline (seed span $0.81$). The masking rate, not the token-level assignment, accounts for most of the measured gain, though the comparison isolating the assignment is inside seed noise.
Chat is not available.
Successful Page Load