Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization
Abstract
Selecting the best large language model (LLM) for a fixed benchmark is often expensive, since exhaustive evaluation requires running every model on every example. Multi-armed bandit (MAB) algorithms can reduce the number of LLM calls by sequentially selecting the next model-example evaluations, thereby avoiding unnecessary computations on clearly underperforming models. Further computational savings can be achieved by predicting the evaluation scores of the various models on the set of examples. In practice, these predictions can be obtained using low-rank (LR) matrix factorization that exploits correlations in the partially observed model–example score matrix. However, such predicted evaluations are not the ground truth: they can be biased and may therefore lead to incorrect identification of the best model. In this work, we propose a principled framework that combines MAB with cheap, predicted evaluation scores without compromising statistical validity. Concretely, leveraging prediction-powered inference (PPI), we derive unbiased estimators for the performance of each model that utilize predicted scores from LR factorization to reduce variance. Importantly, this enables the construction of finite-sample valid confidence intervals in our non-i.i.d. setting, where models are selected adaptively and examples are sampled without replacement. Empirical results on real-world benchmarks demonstrate that our approach reduces the number of required evaluations, leading to meaningful savings in compute and cost, while accurately identifying the best-performing model.