HOP-CAT: Higher-Order Profile-Constrained Adaptive Testing for Efficient Multidomain LLM Evaluation
Abstract
Evaluating large language models on modern benchmarks often requires thousands of item responses, making evaluation costly. Existing efficient evaluation methods typically reduce this burden by selecting small subsets to recover aggregate performance or preserve model rankings. Multidomain benchmarks, however, are intended to support not only an overall comparison but also conclusions about performance across predefined subjects. Preserving the aggregate score alone may therefore fail to preserve the structure and scope of the claims supported by the original benchmark. In practice, accurate aggregate recovery can coexist with substantial subject-level estimation error and uneven subject coverage, with some subjects receiving few or no direct observations. We therefore formulate efficient evaluation on multidomain benchmarks as a coverage-constrained subject-profile estimation problem under a fixed item budget. To address this problem, we introduce HOP-CAT (Higher-Order Profile-Constrained Adaptive Testing), which integrates a higher-order latent-trait model to share information across subjects, an A-optimal item-selection criterion, and a constraint rule that preserves the benchmark blueprint. Across 121,500 end-to-end evaluations spanning six benchmarks and ten random 50/15 model splits, HOP-CAT achieves the lowest cross-budget mean Profile MAE on all six benchmarks and leaves no subject uncovered on any benchmark. Controlled simulations further show that these gains persist across diverse latent structures, calibration sizes, and evaluation budgets. Mechanism analyses indicate that the higher-order ability estimator provides the most consistent improvement. More broadly, this work advances validity-conscious efficient LLM evaluation by preserving the structure of the claims that multidomain benchmarks are intended to support.