Anytime-Valid LLM Evaluation with Cost-Aware Adaptive Generation Depth
Abstract
Evaluating an LLM often involves sampling multiple responses to the same task. Because generations from the same task tend to be clustered, they should not be treated as independent experimental units. More generations reduce uncertainty within tasks, while more tasks provide information across tasks. We study this breadth--depth trade-off in sequential paired model comparison using a bounded task-level aGRAPA e-process with generation depth selected predictably from completed tasks. We relate the classical cost-optimal allocation rule to local e-value growth and derive an effect-aware depth under the quadratic aGRAPA approximation. Across the cost scenarios, the data-driven policies averaged over the six discovery-selected HumanEval comparisons differed from the validation-best fixed depths by at most (1.1) percentage points in rejection probability and (1.5) percentage points of budget in restricted decision cost.