Learning from Disagreement: Maximum Divergence Knowledge Distillation
Aref Jafari ⋅ Parsa Ashrafi Fashi ⋅ Mehdi Rezagholizadeh ⋅ Hanieh Asadi Golmankhaneh ⋅ Shayan Salehi ⋅ Vikram Appia ⋅ Emad Barsoum ⋅ Ali Ghodsi
Abstract
Knowledge distillation pipelines typically train a student to match a teacher on a fixed corpus of teacher-labeled examples. This leaves a major source of supervision unused: the teacher is a fully queryable model whose output distribution extends far beyond any fixed dataset. We argue that effective distillation should exploit this structure by steering generation toward regions where the student most underestimates the teacher. We propose \textit{Maximum Divergence Knowledge Distillation} (MDKD), a rejection-sampling method for autoregressive language models. At each decoding step, MDKD draws candidate tokens from the teacher and preferentially accepts those the student underweights, constructing trajectories that remain teacher-supported while concentrating training on regions where the student assigns insufficient probability. We show that stochastic MDKD's acceptance rule has an exact distributional characterization: it samples precisely from the teacher's \textit{uncovered mass} $[T - S]_+$, the probability the teacher assigns to tokens the student does not yet cover, with stochastic acceptance probability equal to the teacher--student total variation distance. Standard teacher sampling, by contrast, spends most of its updates on the overlap $\min(T, S)$ already absorbed by the student. This makes MDKD the structural complement of speculative decoding, which exploits the same overlap to accelerate inference. On GSM8K, distilling \texttt{Qwen2.5-14B-Instruct} into a base \texttt{Qwen2.5-1.5B} student with only $1{,}000$ KD samples (5 epochs) lifts accuracy from $8.26\%$ to $64.90\%$, outperforming SOTA methods by $10.3$ percentage points and recovering $65\%$ of the teacher--student gap. A divergence-gain analysis confirms the mechanism: MDKD sequences raise student cross-entropy by $43.6\%$ while preserving $98\%$ of teacher-sampling accuracy. Across arithmetic reasoning, dialogue summarization, and code generation, MDKD matches or outperforms strong distillation baselines.
Chat is not available.
Successful Page Load