Conditional Evaluation of Language Models with Cheap Auxiliary Signals
Zhi Zhang ⋅ Lingfeng Lyu ⋅ Yue Kang ⋅ Doudou Zhou
Abstract
Aggregate accuracy hides where models succeed and fail. Estimating conditional performance profiles from gold labels alone is expensive, while cheap auxiliary signals such as LLM-judge scores, pairwise comparisons, confidence scores, and judge-disagreement features can be collected for every benchmark item but are often biased or miscalibrated. We propose \method{} (Local Augmented Control-Variate Evaluation), a semi-supervised estimator for conditional LLM evaluation. The key step is local centering: after subtracting the conditional mean of a cheap signal within the target profile region, any linear augmentation has zero conditional mean and therefore cannot change the estimand. The augmentation coefficient is used only for efficiency, and a local ridge control variate combines a gold-label residual mean from the labeled subset with a cheap-signal mean from the full item pool. We prove calibration-free identification, unbiasedness for grouped profiles, local oracle optimality within centered linear augmentations, and first-order adaptivity to the estimated coefficient. The resulting gain formula depends on a local $R^2$ that can be estimated from the available labeled subset and cheap signals, diagnosing where cheap signals are useful. We extend the same principle to direct paired model gaps and deployment-weighted scores, and validate it on MATH-500, ScienceQA, MMLU-Pro, and GPQA.
Chat is not available.
Successful Page Load