Hard to Find, Easy to Fool: Witness Rank for Budgeted Scientific Verification
Luca Mondonico ⋅ Albert Luo ⋅ Yi Cui ⋅ Zhenan Bao
Abstract
AI scientists and autonomous design systems can rank candidates faster than trusted experiments can evaluate them, yet standard benchmark metrics do not quantify the evidence for the single candidate ultimately nominated. We ask how many trusted evaluations are needed to overturn that nomination. On a declared finite pool, the margin-aware witness rank $R_\Delta$ is the position, under a policy that never reads unobserved trusted labels, of the first candidate that beats the nomination by $\Delta$. Across 230 audits on four fully measured protein and DNA landscapes, including two published offline optimizers, conditional median witness ranks are three to five, and structured policies expose witnesses far more often than random probing. Rank-shuffling locates much of this advantage in the fine ordering of the proxy head. Oracle-controlled rankings with top-tail Spearman correlation near $0.95$ retain conditional medians of two to four; higher margins mainly reduce whether a qualifying witness exists. Retrospective evaluation of 7,300 simulated matched-budget campaigns shows that the same measurements support incumbent-preserving delivery. A confidence-adjusted certificate extends the audit to noisy evaluators, while an independent random channel bounds residual witness prevalence when a noiseless shortlist finds none. Witness rank reports post-selection reliability in experimental-budget units and gives a rule for replacing or retaining a proxy-guided recommendation.
Chat is not available.
Successful Page Load