Human Oversight Is Not a Control Unless the Verification Is Performable
Abstract
Every deployed agentic system whose safety case includes the words "human review" claims that the assigned check can succeed under the conditions furnished. That claim is measurable. Agent benchmarks and measures of the human's behavior – approval rates, review time, overrides – record what the agent can do and what the reviewer did. Behavioral indicators are confounded; a high approval rate can indicate either strong model performance or ceremonial review. This paper argues that evaluation should audit verification conditions directly and should not credit a review step as a functioning control unless the required check is performable before intervention becomes ineffective. It operationalizes a condition-level audit developed within Human, AI, and Organizational Performance (HAOP), a work-system safety framework rooted in occupational-safety practice. At each consequential transition the audit specifies the Verification Requirement, compares it with the Verification Capacity across four conjunctive dimensions – knowledge, evidence, time, and authority – classifies the transition as Covered, Uncovered, or Overrun, and reports the Workflow Profile with raw state counts, the number of transitions audited, and Overrun prevalence, never an average. Repeated audits make human-agent coevolution inspectable as changes in verification requirement, furnished capacity, and transition state. The audit complements model evaluation and behavioral testing. It is not yet empirically validated; the paper sets out a research agenda for testing its reliability and predictive validity.