Efficient evaluation and error pattern discovery for blackbox AI systems
Abstract
Reliable AI evaluation is a prerequisite for measuring progress and for deployment in safety-critical settings. Unfortunately, even gathering frontier system outputs on a full suite of relevant benchmarks can be expensive, let alone adding expert review or running evaluation inside of an optimization loop. Existing approaches to sample-efficient evaluation address this concern only partially: they typically assume abundant calibration data, pretrained error-relevant embeddings, or a fixed benchmark. Meanwhile, optimization loops require repeated evaluation and the ability to surface a long tail of errors; in fact, in real world settings, the long tail of errors often drives human-in-the-loop iterative system improvement. Working in the regime in which the only available information is system inputs and one or more cheap surrogate measures of system error, we derive an optimal importance sampler for error estimation and analyze its theoretical properties. Going further, we show how the same sampler can naturally discover and rank error patterns. On standard benchmarks for frontier model question answering (MMLU-Pro) and agentic coding (TerminalBench 2.0) we show our method consistently outperforms standard Monte Carlo on estimating aggregate error, recovering failure modes and calculating their prevalence, and detecting poison agent/environment injections.