Language Denoising Objectives Extend the Value of Limited Data
Abstract
The choice of language model pretraining objective was once contested among causal language modeling (CLM), masked LM, T5 span corruption, and UL2-style mixtures of denoisers. It was settled in favor of CLM during an era when high-quality text was abundant relative to compute. That era has ended. Compute now routinely outpaces the supply of unique data, and modern pipelines rely heavily on data repetition both during pretraining on curated corpora and during midtraining on small domain corpora. We revisit the objective question in this data-constrained regime. Across a large sweep over models from 15M to 1B parameters, unique data budgets up to 6B tokens, and across repetition budgets, we compare CLM against three denoising objectives: fill-in-the-middle (FIM), T5-style span corruption, and a UL2-style mixture-of-denoisers. All four scale similarly when every training token is unique, but they differ significantly in their tolerance for repeated data. Denoising objectives, which naturally augment language data on every pass, accumulate substantially less overfitting cost than CLM. We find that span corruption and mixture-of-denoisers are the most robust. We then validate the practical payoff in a midtraining setting. We continue web-text pretrained checkpoints on a limited mixture of web-text and math data. We observe denoising midtraining objectives outperform CLM on GSM8k, and the gap widens with increased repetition. Choosing a denoising objective is a simple, complementary lever to model and data scaling for practitioners working under tight data budgets.