Rethinking Knowledge Distillation for Diffusion Language Models
Yuxuan Sun ⋅ Yuanjian Xu ⋅ Jianing Hao ⋅ Yu Li ⋅ Sowmen Das ⋅ Zhong Li ⋅ Sangwoong Yoon ⋅ Miguel Rodrigues
Abstract
Discrete diffusion language models (DLMs) offer a promising non-autoregressive alternative to large language models by enabling parallel generation, yet their performance still lags behind autoregressive counterparts. Knowledge distillation has recently proven highly effective for improving autoregressive models, but its applicability to DLMs remains poorly understood. We present the first systematic study of KD for DLMs and uncover two non-intuitive failures of classical forward-KL (FKL) distillation: (i) student perplexity does not improve monotonically with more teacher-generated data, and (ii) on the same teacher-generated corpus, FKL distillation underperforms simply training the student with the original DLM denoising objective. These failures hold for both a 110M MDLM teacher and a 7B Dream teacher. To address them, we propose a unified, data-free distillation framework for masked DLMs. The framework first runs an off-policy phase that absorbs general knowledge from a static set of teacher samples and then transitions to an on-policy phase whose objective combines: a trajectory loss along the student's own reverse process, confidence reweighting that down-weights tokens where the teacher is unsure, and a trust-region anchor that stabilizes the moving target. For the smaller MDLM regime where FKL alone is insufficient, we additionally adopt the $\alpha$-$\beta$ divergence. On MDLM, the resulting student attains a GPT-2 perplexity of $24.75$, surpassing the teacher's $27.04$; on Dream-7B distillation to a 93M student model, the framework reduces Qwen-2.5 perplexity from the FKL-distilled baseline of $25.23$ to $19.83$ while preserving zero-shot accuracy.
Chat is not available.
Successful Page Load