On the estimation and validity of AI time horizons—a statistical look at the METR plot
Abstract
METR's 50\% time horizon measures the human time length of software tasks an AI solves with 50\% probability, allowing AI capabilities to be put in interpretable units. On 228 tasks and 21 AIs, we recompute the time horizons using splines, as well as item response theory, to relax the assumption that the AI difficulty of a task depends linearly on the log of human time. Our fitted spline can be interpreted as a function which \emph{converts} human time to AI difficulty---and it is nearly flat in a region from 2--15 min, but close to linear elsewhere. Hence, a time horizon jump from 2 min to 15 min is much easier than one from 15 min to 2 hours despite similar multipliers (7.5X vs 8X). Our method improves the time-horizon point estimates on a cross-validated suite of proper scoring rules, though the estimates preserve the overall increasing exponential trend of time horizons in calendar date. We suggest that the time horizons be displayed together with the time-to-difficulty conversion function, especially as new time-horizon-based benchmarks are proposed or existing ones grow to include longer tasks.