Efficient Language Model Evaluation Using Latent-Space Stratification
Abstract
Evaluating a large language model's (LLM) classification accuracy, or other performance metric, over a large corpus can be prohibitively expensive when evaluations are costly and the evaluation budget is limited. We frame LLM evaluation as a sampling problem: given a large candidate pool, the goal is to estimate overall performance as accurately as possible under a fixed evaluation budget. We study this setting in the presence of latent structure derived from document embeddings. Our method begins with a small random sample, fits a Gaussian process classifier of correctness over the latent space, and uses the predicted correctness surface to define strata for second-stage sampling. The remaining evaluation budget is then distributed across strata using a Neyman-style rule based on stratum size and estimated variability, and overall performance is estimated using a standard stratified estimator. We evaluate the approach on synthetic data and a real arXiv document classification task. In both settings, the method yields lower-variance estimates than simple random sampling under the same budget when the latent representation is informative of LLM correctness. More broadly, these results demonstrate the potential of model-assisted sampling to improve the efficiency of LLM evaluation when exhaustive labeling is prohibitively expensive.