Efficient Sequential Evaluation of Large Language Models
Chia-Yu Hsu ⋅ Shubhanshu Shekhar
Abstract
We study the problem of estimating the average correctness of a new large language model (LLM) over a fixed question bank of $N$ questions. Prior work has mainly considered non-sequential settings, which require choosing the number of queried questions in advance and may waste evaluation resources. We propose a sequential evaluation framework that allows stopping once the desired estimation accuracy is reached while maintaining anytime-valid statistical guarantees. To improve efficiency, we actively select questions to shrink the uncertainty as quickly as possible. Technically, we leverage question-level correctness predictions obtained from historical data to construct a confidence sequence (CS) using reverse information projection (RIPr), propose a max-min querying rule based on predicted one-step log-growth, and analyze how prediction mismatch affects the CS shrinkage rate.
Chat is not available.
Successful Page Load