Speculative Evaluation of Stochastic LLMs
Abstract
Evaluating a stochastic large language model is costly: benchmark scores estimate expected performance from randomized rollouts, yet uniform repetition ignores sharp differences in task-level rollout variance. We ask how to minimize the variance of a fixed-benchmark mean under an exact rollout budget. We develop Speculative Evaluation, which runs a short uniform pilot, pools per-task success counts with a hierarchical Bayesian model, and uses posterior expectations of task-level sampling variances for exact positive-integer Neyman allocation. Fresh, pilot-independent continuation samples and deterministic stage weights yield an exactly unbiased estimator for every fixed task-probability profile, regardless of prior misspecification. Across six checkpoints and 18 benchmark groups, we evaluate 107 nondegenerate benchmark–checkpoint pairs. For average rollout budgets from 8 to 64, Speculative Evaluation reduces mean variance by 12.8%–33.6% relative to uniform evaluation.