Calibration without Ground Truth
Yuqing Kong ⋅ Mingyu Song ⋅ Yizhou Wang ⋅ Yifan Wu
Abstract
We propose a label-free post-processing framework that improves a strong but miscalibrated primary model using a weaker yet better-calibrated reference. Our key insight is that strict improvement is possible if and only if the two models are not *mutually calibrated*, meaning there does not exist joint distribution over their predictions and outcomes such that both predictions are simultaneously calibrated. We formalize this condition and connect it to no-arbitrage results from economics. Under such condition, we develop an efficient post-processing algorithm of the strong model's outputs based on Bregman projection, with a strict worst-case performance improvement guarantee. Experiments on representative LLMs across varying scales demonstrate the effectiveness of our method, reducing the ECE of the primary model by over $40$\% on common benchmarks.
Chat is not available.
Successful Page Load