Training Optimal Large Diffusion Language Models
Jinjie Ni ⋅ Qian Liu ⋅ Chao Du ⋅ Longxu Dou ⋅ Hang Yan ⋅ Zili Wang ⋅ Tianyu Pang ⋅ Michael Shieh
Abstract
We introduce Quokka, the first comprehensive scaling law for diffusion language models (DLMs), encompassing both compute- and data-constrained regimes alongside key modeling and optimization designs. Under compute constraints, we find that optimal parameter and dataset sizes scale equally with compute ($N_{\mathrm{opt}} \propto C^{0.5}$, $D_{\mathrm{opt}} \propto C^{0.5}$); however, DLMs are 2--5$\times$ more data-hungry than autoregressive (AR) models, requiring larger corpora and proportionally smaller models for a given FLOP budget. Under data constraints, validation loss follows a U-shaped curve across epochs, with the onset of overfitting scaling as $e_{\mathrm{opt}} \propto U_D^{0.39}/N^{0.55}$. This indicates that optimally leveraging a fixed unique data budget ($U_D$) requires allocating both modestly larger models and more training epochs. Beyond data and parameter allocation, we provide actionable guidance on DLM design choices: absorbing-mask transitions outperform uniform diffusion, linear noise schedules are the most stable and performant, and principled diffusion ELBO objectives ultimately surpass MaskGIT losses despite slower initial convergence. Furthermore, we demonstrate that easy-to-hard noise curricula accelerate early learning, established AR scaling laws for batch size and learning rate transfer seamlessly to DLMs, and weight decay, while unhelpful for single-epoch runs, is essential for multi-epoch regularization.
Chat is not available.
Successful Page Load