Truthful Calibration Errors for Multi-Class Prediction
Abstract
Calibrated predictions are useful because their numerical values can be interpreted as probabilities. Calibration errors are therefore widely used to evaluate and compare probabilistic predictors. Recently, Haghtalab et al. [2024] introduced truthfulness as an additional requirement for such measures. A calibration measure is truthful if a predictor minimizes its expected measured error by reporting the true conditional label distribution. Many standard empirical calibration errors are non-truthful: a predictor may appear better calibrated by distorting its probabilities. We study the implications of truthful and non-truthful calibration errors for practice. First, we introduce perfectly truthful calibration errors for full multiclass calibration and classwise calibration, two standard multiclass notions. More generally, our construction applies to any linear property of the label distribution, generalizing the truthful calibration error for binary predictions in Hartline et al. [2025]. We also identify a truthful correction for confidence calibration. Second, we characterize the decision-theoretic implications of these truthful errors. For calibrated predictors, truthful calibration errors preserve the Blackwell dominance: a more informative calibrated predictor receives no larger expected error. Third, we show that this decision-theoretic interpretation explains and mitigates the well-observed ranking robustness problem of binned calibration errors. Empirically, non-truthful confidence-based errors can reverse model rankings when the number of bins changes, while our truthful classwise error gives more stable rankings across binning choices.