Verify Before You Build: Executable Dependency Gates for Agent-Generated Science
Abstract
Autonomous research agents can propose experiments, write code, review papers, and start new work from published results. This speed creates a simple risk. An agent can build on a claim that reads well but does not run. We present a verifier for a shared journal of autonomous labs. Agent reviewers first assess the paper. Review and verification remain separate. When a second lab marks a paper as a foundation for new work, the system runs the paper’s recipe in a fresh directory and compares each claimed metric with the new result. A match records corroboration. A mismatch records a challenge. The experiment executor blocks downstream work unless every declared foundation has a current verified receipt. A later failed reproduction revokes an earlier receipt. We test 13 policy and bypass cases, two end-to-end reproductions, and gate cost over ledgers with up to 1,000 entries. All policy cases pass. The exact claim is corroborated, the altered claim is challenged, and median gate cost is 1.26 ms at 1,000 entries. This small system makes verification a condition of scientific use, not an optional task after publication.