CELEUS: Certifiable and Efficient LLM Evaluation via E-Processes
Abstract
Can we trust evaluation scores to capture an LLM's true real-world performance? Certifiable evaluation addresses this question by providing statistical guarantees for LLM evaluation. In particular, existing methods sequentially select evaluation samples and update confidence intervals (CIs) intended to cover the true performance with high probability (e.g., 95%) until a stopping criterion is met, such as the CI width reaching a target precision. However, existing methods are not generally anytime-valid: the claimed coverage may fail when CIs are repeatedly updated and used to determine when to stop, leaving a gap between theoretical rigor and practice. This paper bridges this gap by proposing CELEUS, a Certifiable framework for Efficient LLM evaluation that leverages E-processes to construct anytime-valid CIs. Concretely, we propose a new inferential signal that combines two ingredients: (i) Uncertainty-guided sampling to select informative samples for evaluation and (ii) Surrogate-assisted approximation to provide approximate scores for unevaluated samples. We prove that this signal remains conditionally unbiased for the evaluation score given the observed history, enabling statistically grounded and anytime-valid e-process CIs. More importantly, these two ingredients reduce estimation variance and help reach the target precision with fewer evaluated samples. We further prove that the widths of the resulting CIs shrink at a near-parametric rate up to logarithmic factors and characterize an oracle variance-optimal sampling rule that motivates our practical uncertainty-guided strategy. Experiments show that CELEUS reaches the target precision using 54–62% fewer evaluated samples than the baselines while preserving anytime-valid coverage.