Claim-Gating under Gated Medical Data: A Barrett Verification Case Study
Abstract
Surrogate experiments can aid engineering while failing to verify the target claim that motivated them. We present an executable claim gate for AI-scientist workflows. A claim declares its statement, full evaluation context, and nonempty typed requirements. A verifier counts only if its question and scope match, its claim-definition digest binds the same statement and context, its required artifacts are byte-consistent, and a recognized content validator passes. The gate emits LICENSED, STOP, or PIVOT, plus the first unsatisfied question. We apply it to a frozen Barrett-neoplasia repository snapshot whose original goal was low-prevalence detection on RARE25. An executable inventory replays 44 unique result-path declarations (36 file-like and eight directory-like) against a sanitized snapshot manifest; none is present, and the snapshot also lacks RARE images and held-out target output. The original superiority claim therefore pivots to two licensed narrower claims. Twenty-eight encoded scenarios---25 rejection attacks, two licensed controls, and the case-study pivot---test vacuous schemas, scope and context substitution, missing evidence, corrupt hashes, rehashed inconsistent content, and failed dependencies. As a worked analytic verifier, we show that at 90% recall and 1000 normal cases per positive, PPV 0.50 requires FPR at most 0.0009; observing zero false positives needs at least 3,328 independent representative normal cases for a one-sided 95% exact upper bound to meet that requirement. These are documentary and analytic decisions, not estimates of model performance.