Beyond Task Success: A Validity-Centered Position on Evaluating Human-Agent Teams in Deployment
Abstract
Agent evaluations increasingly summarize system quality with task completion, accuracy, preference scores, or aggregate judge ratings. We argue that these measurements are useful but insufficient: a score is evidence about a system, not a guarantee that the evaluative claim attached to that score is valid. This distinction becomes consequential for deployed human--agent systems, where goals are negotiated, human feedback is heterogeneous, workflows constrain what actions are useful or safe, and both users and agents change over time. We advance a validity-centered position: evaluation should be organized around the claims that stakeholders intend to make, the evidence used to support those claims, and the deployment scope within which the inference is justified. We develop five propositions concerning construct specification, deployment transfer, human disagreement, interaction effects, and temporal validity. We then propose a practical claim--evidence--scope protocol for benchmark design and reporting. The protocol does not replace performance metrics; it constrains what those metrics are permitted to mean. We close with a research agenda for disagreement-preserving evaluation, longitudinal validity audits, trajectory-level assessment, and validity-aware benchmark reporting. The goal is to shift agent evaluation from "what score did the system obtain?'' toward "what conclusion does this evidence actually justify about a human--agent team in use?''