From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation
Bora Kargi ⋅ David Salinas
Abstract
Evaluating new large language models typically requires costly human annotation campaigns at scale. LLM-as-a-judge offers a cheaper alternative, but judge scores carry systematic errors — such as position bias, self-preference, or intransitivity — that can strongly miscalibrate the resulting rankings. We quantify the resulting judge--human disagreement at two complementary levels. At the local level, we estimate per-battle uncertainty from the judge's own score differences by propagating calibrated win probabilities rather than hard labels into the Bradley--Terry procedure. This alone provides a drastic improvement to Elo estimation accuracy, bringing LLM-derived ratings within $17.9$ Elo MAE of human-derived ones when averaged over 55 held-out models on LMArena. At the global level, we apply split conformal prediction to the residual gap between LLM-derived and human-derived Elo ratings across held-out models, producing prediction intervals with distribution-free marginal coverage guarantees that account for irreducible LLM--human disagreement. Together, these two layers yield a low-cost evaluation tool that provides developers with calibrated Elo estimates and honest uncertainty bounds, without access to large-scale human annotations.
Chat is not available.
Successful Page Load