Joint Sequence--Vocabulary Selection for Efficient LLM Distillation
Xueli Geng ⋅ Weicheng Zhao ⋅ Xutong Mu ⋅ Tianrui Wei ⋅ Yanbiao Ma ⋅ Yulong Shen
Abstract
Selective knowledge distillation has become increasingly important for compressing large language models, where dense logit-based knowledge transfer over long sequences and large vocabularies incurs substantial computational overhead. Distillation supervision is naturally organized over a sequence–vocabulary matrix, with its highest-information components often concentrated on a small subset of specific sequence–vocabulary pairs. Existing methods typically select sequence positions or vocabulary candidates independently, relying on coarse-grained one-dimensional heuristics that weaken the precision and effectiveness of knowledge transfer. To address this, we propose $\textbf{J}$oint $\textbf{S}$equence–$\textbf{V}$ocabulary $\textbf{D}$istillation ($\textbf{JSVD}$), a plug-and-play selective distillation framework that identifies and focuses supervision on informative sequence–vocabulary pairs. For each training example, JSVD constructs bidirectional teacher--student candidate supports, scores sequence--vocabulary pairs using prediction discrepancies, and progressively focuses distillation on unresolved informative pairs during training. Extensive experiments across diverse distillation objectives, model pairs, and benchmarks show that JSVD consistently improves downstream performance while reducing distillation overhead and accelerating convergence.
Chat is not available.
Successful Page Load