Task Informativeness in AI Evaluation Is Objective-Dependent
Victoria Kurakata
Abstract
AI capability evaluations rely on a finite set of tasks to inform high-stakes decisions. Evaluations may fulfill at least two distinct objectives: they may provide a measure of scalar capability or predict performance on unevaluated tasks. We provide theoretical and empirical evidence to suggest that the tasks most useful for one objective are not necessarily the most useful for the other: a task is equally informative for the two objectives only when model performance is effectively one-dimensional. Using run-level data from METR's Time Horizon 1.1, we use held-out model folds to measure the strength of random and objective-selected suites. For random suites, the rank correlation between measurement loss and prediction loss is between $-0.02$ and 0.05, indicating that a suite's informative value for one objective is virtually uncorrelated with its value for another objective. However, certain suites are indeed more informative than others: suites selected for a stated objective reduce held-out error by 32–60\% for measurement and by 17–23\% for prediction. Selection for one objective also degrades performance on the other: a suite chosen for measurement raises omitted-task prediction error by 13–65\%, and a suite chosen for prediction misses the full-suite time horizon by a factor of 2.5–4.6. We conclude that evaluation suites ought to be constructed and audited relative to the particular objective they are intended to support.
Chat is not available.
Successful Page Load