Benchmarks as Measurement Instruments: Quantifying Signal and Noise for More Efficient AI Evaluations Under Distribution Shift
Michael Hardy ⋅ Anka Reuel-Lamparth ⋅ Jodi Casabianca ⋅ Hansol Lee ⋅ Benjamin Domingue ⋅ Sanmi Koyejo ⋅ Mykel J Kochenderfer
Abstract
AI progress is increasingly tracked through evaluations, such as AI benchmarks, yet small perturbations in evaluation pipelines can produce unstable scores and model rankings. We recast benchmark reliability as a measurement problem and formalize it as a signal-to-noise ratio derived from a crossed random-effects decomposition. This framework yields a direct mapping between variance components and expected rank stability, linking theoretical reliability to observable leaderboard concordance. Modern AI benchmarks operate in a high-dimensional regime with many items and relatively few evaluated models, where classical item-level reliability measures are ill-posed. We define a lower bound for benchmark item-level reliability classically, and we further introduce a tractable proxy, $\lambda^\bigstar_6$, that preserves the variance structure underlying reliability without modifying the estimand through sparsity or dimensionality reduction. Across diverse benchmarks, we show that selecting items via $\lambda^\bigstar_6$ produces smaller subsets that achieve higher ranking reliability than random subsampling at fixed evaluation budgets. Finally, we analyze reliability under positive distribution shift--as often observed in the current AI ecosystem--by partitioning models into lower- and higher-performing cohorts. We show that reliability is population-dependent and frequently declines as between-model variance shifts among stronger systems. Our results establish a principled framework for diagnosing and improving benchmark reliability and demonstrate that ranking stability is neither intrinsic nor static, but a measurable and optimizable property of the benchmark–population pairing.
Chat is not available.
Successful Page Load