Count Every Validation: Auditing Hidden Attempts in Recursive AI Scientists
Abstract
AI scientists may open and suppress more outcome-bearing validations than appear in the hypothesis stream seen by an online error controller. We introduce DiscoveryStream Audit, a broker-mediated protocol that binds confirmatory attempts before access and reconciles every receipt to a controller decision and typed terminal release. In frozen held-out global-null experiments, winner-only suppression increased e-LOND false-release probability relative to complete batch registration by 0.0293 (95% paired interval 0.0270–0.0316) under correlated best-of-16 selection and by 0.00345 (0.00264–0.00426) under directional-sign selection; the conservative e-LOND schedule remained below 0.05, while Bonferroni error reached 0.4996 versus 0.0487 in the correlated case. In a depth-six recursive forest, suppression opened 4,000 outcomes behind 500 registered positions and yielded Bonferroni FWER 0.2670 versus 0.0036 with complete registration. Replaying sampled records through the auditor made all 10 suppressed traces fail and all 30 sealed, prebound, or fresh-retest traces pass; a generated conformance suite and a two-attempt pinned integration exercised the same lifecycle boundary. Valid routes triggered no prespecified simultaneous calibration failure, but some were extremely low-power. The contribution is broker-relative attempt accounting for existing controllers—not a new FDR theorem or a guarantee of scientific truth.