Before Measuring Human-Agent Teams: Validating Human References in Agent Benchmarks
Anqi Peter Li ⋅ Ethan Yip ⋅ Kundana C Kommini
Abstract
Agent benchmarks increasingly report capability on a human scale: the length of a task a model can finish. That scale now also carries a claim about the rate of progress and about what work agents can reliably take over. Every such claim divides by a human number that the benchmark measured for itself, yet later analyses treat that number as known. We audit this number, the referent, across five public agent benchmarks. Its reliability is the share of a released difficulty axis that is signal; the rest is measurement error. We estimate it from data the releases already contain: repeated timings, the width of a released band, or an external human-calibrated scale. On SWE-bench Verified, most of the axis's variance is noise. When we carry reliability through the capability estimator, the usual correction does not transfer: each route implies a different observation model, and only repeated timings license an attenuation rescale. Our second correction, for how each value was obtained, is weaker. The matched contrast between timed and elicited tasks moves the reported doubling time by under $7.9$ days over paired resamples, which is compatible with zero. A joint latent model gives a much larger model-conditional shift, but that shift is not data-identified, so the data cannot adjudicate between the two. We therefore argue that any benchmark read as a claim about human substitution should publish three numbers: (a) the reliability of its referent, (b) how its values were obtained, and (c) its timed support at the operating point. More broadly, practitioners should take care before they treat the human side of an agent benchmark as a number they already know. This audit covers that measurement only and does not estimate human-agent team productivity.
Chat is not available.
Successful Page Load