The Wrong Data at the Right Time: Off-Domain Mixing for Discrete Diffusion
Abstract
We show how to use out-of-distribution text to improve the quality of discrete diffusion models. Typically, such models are trained on a single curated corpus, and when the target domain is small, related corpora are either discarded or naively mixed in, which biases generation away from the target. We show that there is substantial value in this related data if it is used at the right diffusion times. We present RefineMix, a simple, principled framework that admits off-domain data only at low noise levels and reserves high noise levels for target data alone. Our framework exploits one property of text under token corruption: a few surviving domain-specific tokens reveal the domain of a sequence even when most of it is masked. We validate the framework on domain-shifted language modeling, where RefineMix improves target perplexity over fine-tuning and pooling on four of five domain pairs, gains more as the target corpus shrinks, and keeps its samples in the target domain where pooling drifts to the source. The core insight is that, unlike Gaussian noise on images, masking does not erase the domain of a text, so out-of-distribution data belongs at low noise. We provide theoretical justification for our approach by bounding the probability that the sampler leaves the target domain and by analyzing the trade-off between biased off-domain data and limited target data across diffusion times.