BoundKD: Distillation as Constraint Satisfaction under Sampled Supervision
Abstract
We study black-box distillation: training a compact student from a teacher that returns only samples, never logits. From N samples per context, we compile certified lower bounds on the teacher's pairwise token preferences, train the student to satisfy each bound as a constraint (bound-constrained knowledge distillation, BoundKD), and reuse the same constraints after training as an audit. Across six teacher-student pairs at 3× to 98× compression, BoundKD improves every ranking metric in every run on a 98× smaller student (Pythia-6.9B to 70M), ahead of sampled-KL throughout, and never raises held-out perplexity more than 2\% above baseline at the smallest budget, while regressing onto the same bounds with Margin-MSE inflates perplexity up to 443×. A certification-ceiling lemma explains why: no finite budget can certify a large margin, so the bounds systematically understate the teacher's confident preferences, and regression trains the student to reproduce that error.