How Much Can Repeated Running Buy? A Measured Variance Ceiling for LLM Code Evaluation
Yongxi Zhou ⋅ Lai Y Choi ⋅ Jiaxi Wen ⋅ Wenbo Ye ⋅ Junwei Yao
Abstract
A benchmark score is a measurement whose uncertainty has two sources: an aleatoric component from stochastic decoding and an epistemic component from the finite problem set. Using 16 models on 99 recent LeetCode-style problems with two prompts and five runs each (15,995 judged generations), we estimate both components and ask a design question: how much can repeated running buy? First, the pooled-binomial interval that repeated-run protocols invite answers neither question of interest---it is $2.0\times$ too narrow for generalisation to new problems and $1.9\times$ too wide for the fixed problem set---and under correctly clustered intervals, 11 of 15 adjacent leaderboard pairs are statistically unresolved. Second, the variance ceiling $1+\mu_w/\sigma_b^2$, the largest factor by which repetition can reduce score variance, has median 1.34 across models (range 1.05--1.71), so repetition can narrow confidence intervals by at most 14--23% on this benchmark. Third, once the cost of adding verified problems is priced in, the precision-optimal repeat count is $R^{*}=\sqrt{\mu_w c_p/(\sigma_b^2 c_g)}$, and the variance curve is flat around it: the decisive reason to repeat is not precision but identification---two to three runs suffice to estimate the components that correct inference requires. We state explicitly what these single-benchmark, single-temperature measurements do and do not license.
Chat is not available.
Successful Page Load