Principled Evaluation Dataset Sizing for LLM Applications
Abstract
How much data is enough to support evaluation of an LLM application? Existing guidance on evaluation dataset size remains fragmented and often relies on rules of thumb that do not specify the precision or statistical power the resulting evaluation can support. The question is sharpest in clinical settings, where ground truth requires expert clinician time and patient data collected under consent. We develop a statistical framework for sizing datasets for offline, system-level evaluation of generative text quality. We distinguish benchmark planning, where data are collected before the prompts or system variants to be evaluated are known, from paired comparison, where two fixed alternatives are evaluated on the same examples. We show how variance estimates from a pilot or reference dataset can be translated into sample-size requirements under either objective. Using an evaluation of LLM-generated summaries for 2,000 CNN/DailyMail articles, we further separate variation between evaluation examples from stochastic variation introduced by repeated LLM generation and judging. Repeated inference can reduce run-level uncertainty and therefore the number of labeled examples required, particularly for paired comparisons, but only to a floor determined by variation between distinct examples. The resulting calculations provide a statistical minimum rather than a universal dataset-size recommendation. Representativeness, data slices, prompt optimization, application risk, and deployment scope may require additional evidence. This framework provides a principled basis for designing evaluation datasets that are both statistically defensible and proportionate to the cost of obtaining ground truth.