Oversight Is a Schedulable Resource: Fail-Closed Scheduling and a Closed-Loop Negative Result
Abstract
Trusted oversight in agentic systems is finite: reviewers, high-assurance models, simulators, and verifiers can all become contended services. We make oversight an AgenticOS resource by representing uncertain decisions as jobs with tenant identity, deadline, uncertainty set, and a declared fallback. A scheduler decides which jobs receive fresh review, but a separate runtime permits exactly one outcome: a state-bound trusted response, the verified fallback, or abstention. We implement this contract with bounded admission, ten scheduling policies, a calibrated value model, and a hash-chained ledger. Fresh Sailboat and PointGoal experiments produce 1,533 locked-test jobs across 60 episode/layout components, 20 audit-corrected stress blocks, and 40 fresh positive-slack blocks. On the fixed trace workload, learned value exceeds earliest-deadline-first (EDF) by 0.077 opportunity units per component on average, with order-specific differences from -0.019 to 0.120. In the fresh-seed replication, learned-minus-EDF is -11.259 task-utility units per block (95% CI [-33.075, 10.768], sign-flip p = 0.341), providing no evidence of a learned advantage. Learned value is also worse than always fallback (−52.63, Holm-adjusted p = 0.0045). All emitted actions satisfy the runtime provenance contract. The result is a caution for AgenticOS design: scheduling makes overload explicit and fail-closed, but a plausible local value proxy need not improve downstream utility.