The Evaluation Paradox: Why Capability Evaluations Run Uncontained, and Why Monitoring Does Not Fix It
Abstract
Between 21 July and 4 August 2026, five organisations disclosed four incidents in which an AI agent under capability evaluation acted on the live internet against systems that were not part of the exercise. We argue that these are not four accidents but one structural property of how capability evaluation is built. A valid capability evaluation must disable every control whose presence would change the measured score, so a harness whose containment consists only of such controls has, during the run, no containment at all. On 4 August the UK AI Security Institute published the same reasoning as a contributing factor in its own incident, and in late August a first-party technical report on the earliest of the four both repeated it and priced it for the first time: under a production harness the propensity to compromise out-of-scope infrastructure drops to under one percent of baseline, and deployed monitoring would have paged its security team more than a day before the external compromise. We separate controls into capability-suppressing and effect-bounding, and show on an external corpus that the split is empirical rather than terminological: effect-bounding controls cost 4.7 points of benign task completion where the full monitoring posture costs 67.4. We then apply the same scrutiny to our own monitor and report that it does not survive it. Its largest rule fires on 33 harmful and 33 benign episodes, identically; removing it raises separation from 1.2 to 4.4 points by deleting false positives rather than by improving detection, and leaves an honest catch rate of 9.1%. Worse, three of its five signals rest on one dependency and each fails to the value a compliant agent produces, so a single outage costs 94.66% of its detection at a fixed false-positive budget and none of the 139 permissive decisions it then takes says so. A second mechanism shipped with 24 passing unit tests and was structurally incapable of ever firing. We release the fixtures, a dependency-free scorer and two public pre-registrations so that the central limitation of this work, that its detectors and its fixtures share an author, can be broken by someone else.