Beyond Agent Accuracy: Evaluating Defensible Work Products in Human–AI Teams
Abstract
Agent evaluations commonly measure whether a system produces a correct answer or completes a task. In consequential institutional workflows, those measures can miss defects in the work product that an institution may later record or rely upon: the cited rule may not apply, evidence may be insufficient, a material limitation may be hidden by a binary answer, or a conclusion may be impossible for another qualified reviewer to reconstruct. We propose a protocol for evaluating human–AI teams by the independently adjudicated defensibility of their work products rather than by answer accuracy alone. The protocol treats anti-circularity as a design requirement rather than an achieved property: empirical success criteria must be derived from external domain norms, frozen before system outputs are observed, and applied by adjudicators who do not use architecture-internal conformance as evidence of success. For a bounded U.S. public-sector audit setting, we provide an illustrative seed rubric traced to the Government Auditing Standards and specify three conditions: model-only, conventional human–AI review, and consequence-aware human–AI review. The load-bearing contrast is the differential between the two human–AI conditions, both of which seek audit quality. Primary outcomes are domain-adjudicated work-product acceptability and unsupported-conclusion rate. Secondary measures include not-established handling, independent reconstruction, adjudicator disagreement, authorization-transition compliance, reviewer burden, and efficiency. Four mechanisms are examined through a staged ablation ladder that makes their couplings explicit. We do not claim that “consequence” is a new concept, that the seed rubric itself proves criterion independence, that the protocol establishes institutional legitimacy, or that the reference workflow is production-ready. The narrower contribution is an executable evaluation protocol whose claims can fail under independent domain adjudication.