Integrity Is Not Validity: A Layered Assurance Case for Stateful Agent Evaluation
Abstract
A reproducible agent experiment can still answer the wrong question. State restoration may be exact while an adapter truncates later actions; event ledgers may reconcile while a scorer enforces an unstated serialisation rule; final states may be correct while agents omit a required verification step. We present a layered audit framework that separates five claims usually collapsed into one benchmark score: adapter reliability, artifact integrity, evaluator-contract validity, objective correctness, and verified workflow capability. The framework uses canonical checkpoints, complete randomised continuation blocks, event-derived outcomes, content locks, and external revision boundaries. We illustrate the layers through preserved development failures rather than filtering them away. A single-request calibration missed later-step truncation in 14 of 16 stateful forks. A subsequent run was internally valid but encoded an unintended byte-equality requirement and a vacuous success condition. After a preregistered semantic repair, all 16 forks reconciled with zero protocol failures and 14 final workspaces were semantically correct, yet only four forks invoked and passed the required evaluator, and only one of four templates met the frozen capability criterion. The registered decision was Defer. In a separate post-registered validator study, all 12 specified corruptions were detected across eight integrity layers; nine rebuilt the outer run lock and three of those also refreshed dependent inner commitments. We show why two executions of one scorer are redundancy rather than independent semantic evaluation. A prospective post-V3 design simulation finds that 96 families is the first tested count meeting interval-width and assumed-equivalence decision-rate targets, but no candidate through 256 passes an uncertainty-aware boundary-error screen. Confirmatory sizing therefore remains unresolved. The contribution is a method for localising invalidity, not evidence for routing value, model superiority, or benchmark portability.