Calibeating Prediction-Powered Inference
Abstract
Prediction-powered inference uses black-box predictions to improve estimation from small labeled samples and large unlabeled samples, but raw prediction scores are often miscalibrated and can be inefficient regression adjustments. We study semisupervised mean estimation with a black-box score, a small labeled sample, and a large unlabeled sample. AIPW and PPI give valid inference for fixed or cross-fitted scores, but their efficiency depends on how well the score predicts the outcome. We propose \emph{calibrated prediction-powered inference}: post-hoc calibrate the score on the labeled sample, then average the calibrated predictions over the pooled covariate sample. The estimator requires no retraining, takes a simple plug-in form, and has an exact AIPW representation. For linear calibration, we show first-order equivalence to PPI++. For isotonic calibration, we establish asymptotic normality, valid Wald inference, and ``calibeating" guarantees: isotonic post-processing improves the score as a predictor and as a first-order regression adjustment within the monotone class, and subsequent score-only post-processing yields no additional first-order gain. We also show that the original PPI estimator is a special case of AIPW and can be inefficient when the prediction score is already accurate. Simulations, benchmark reproductions, and an LLM-evaluation application show that calibrated estimators often improve on PPI and are competitive with AIPW and PPI++.