What Does Uniform Noise Add Beyond MASK?
Abstract
Two paradigms currently dominate the discrete diffusion landscape: masked and uniform diffusion. In masked diffusion, tokens are iteratively replaced by masks in the forward process; while in uniform, they are replaced by tokens drawn uniformly from the vocabulary. While these processes may look unrelated, can we connect their learning problems? In this paper, we first decompose the continuous-time training objective of these processes into two terms: a conditional-replacement loss and an exit-rate loss. The exit rate controls how quickly a token is replaced and is directly prescribed by the noise schedule in masked diffusion. Under uniform diffusion, we show that the exit-rate loss trains the model to estimate whether the current token matches its clean source; at a given noise level, lower confidence leads to faster replacement. Empirically, we find that uniform replacements rarely pass for contextually plausible alternatives to the clean tokens. These findings motivate semantically structured noising processes that preserve token revisability while providing more frequent supervision for detecting and correcting plausible mistakes.