PACE: A Proxy for Agentic Capability Evaluation
Yueqi Song ⋅ Lintang Sutawika ⋅ Jiarui Liu ⋅ Lindia Tjuatja ⋅ Jiayi Geng ⋅ Yunze Xiao ⋅ Daniel Lee ⋅ Aditya B Soni ⋅ Vincent Lo ⋅ Xiang Yue ⋅ Graham Neubig
Abstract
\begin{abstract} Evaluating large language model (LLM) agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to complete. In contrast, non-agentic LLM benchmarks that test individual capabilities (e.g., reasoning, code generation, instruction following) are fast and cheap to run. We investigate whether performance on expensive agentic benchmarks can be accurately predicted by the performance on a small, carefully selected subset of atomic evaluation instances. We introduce PACE, a framework that constructs proxy benchmarks by selecting instances from existing non-agentic evaluations whose aggregate scores most reliably predict agentic benchmark performance. Given a pool of candidate instances spanning atomic capabilities (instruction following, planning, tool calling, etc.), PACE fits a regression that maps a model's scores on a compact subset of source instances to its score on the target agentic benchmark; the subset itself is produced by combining two complementary instance-selection strategies, a target-relevance local selection and a globally informative global selection. Experiments across 14 models, 4 agentic benchmarks, and 19 non-agentic benchmarks show that PACE-Bench predicts agentic scores with leave-one-model-out (LOOCV) mean absolute error (MAE) under $4\%$, Spearman correlation above $0.80$, and pairwise model-ranking accuracy around $85\%$, all at much less than $1\%$ of the full agentic evaluation cost. We further analyze the selected proxy instances, revealing what skills each agentic benchmark uniquely demands. PACE enables practitioners to obtain reliable estimates of agentic performance during model development, selection, and routing, without the overhead of full agent evaluation.
Chat is not available.
Successful Page Load