Empirical Verificational Complexity: A Research Programme for Scalable Oversight and Steering in Autonomous Laboratories
Abstract
Verification of machine-generated science is usually attempted over the published record. We take it to belong in the autonomous laboratory, where machine-generated results are produced and where the measurement that would settle a claim can be performed, and measure one instance of it. This paper defines a verification contract, which fixes what counts as having checked a claim, and builds a benchmark on a published dispute over an autonomous laboratory's synthesis claims: 95 targets drawn across the original claims, an independent critique, and the correction that adjudicated each one. The benchmark tests whether a verifier, given only the deposited evidence the expert panels held, independently reaches the findings they published. One architecture satisfies the contract: a model proposes checks and rival hypotheses, and every verdict is re-established by a small independent, trusted kernel. This verification layer reaches 55 of 55 stage-one targets; the best unlayered condition reaches 41. No verifier reproduces the published adjudication from the evidence available when the claims were made, including the system the benchmark scores against. Two properties follow. The layer produces the data steering consumes, in that what it asks for next depends on the evidence it already holds. And oversight is scalable, in that the component that must be trusted for a refutation to count grows with the variety of the record rather than its volume. The measurements ground a programme: what a claim costs to audit is a measurable quantity, and the laboratory is where it is measured.