Prediction-only distillation with optimal mixing in ridge-regularized linear and logistic regression
Hien Dang ⋅ Pratik Patil ⋅ Alessandro Rinaldo
Abstract
Self-distillation is typically studied when the student is retrained on the teacher’s original training inputs. In many realistic deployments, however, the labeled data are unavailable after training, and one only has access to the trained predictor and fresh unlabeled covariates. We study distillation in this prediction-only regime through a fresh-$X$ prediction-mixing scheme: a pure-distilled student is trained on pseudo-labeled fresh features, and the final predictor is formed by an affine combination of the teacher and student predictions with the optimal unconstrained mixing weight. For ridge regression under proportional asymptotics, we derive deterministic equivalents for the optimally mixed risk under general anisotropic covariance and deterministic signal, and show that optimal mixing strictly improves upon the teacher for almost every pair of regularization levels. In the isotropic case, we uncover a sharp and counterintuitive phenomenon: the optimally mixed risk is generically non-monotone in the amount of unlabeled data. To tune the optimal mixing weight, we show that a small independent labeled calibration set suffices for consistent one-shot estimation at low computational cost, since no additional retraining is required, whereas such tuning is impossible when using only teacher’s outputs on unlabeled data. Finally, to extend beyond squared loss, we analyze logistic regression for binary classification and show that prediction mixing can improve over both the teacher and the pure-distilled classifier.
Chat is not available.
Successful Page Load