s the Endpoint Well-Posed? Measurement Contracts for AI-Assisted Scientific Discovery
Abstract
Claims about AI-assisted scientific discovery increasingly rest on benchmarks that reward agreement with a single historical paper endpoint. Before such scores are interpreted as evidence of hypothesis formation or scientific competence, the benchmark's measurement contract must itself be validated. We show that this contract can fail in two directions at once: inputs may over-determine the endpoint (availability Ravail) or under-determine it (identifiability Rid). We introduce a dual-axis measurement-contract framework using scorer-disjoint measures, falsifiers, model-free baselines, and length-nearest evidence-identity interventions. Empirically, on a CS/AI-for-science disclosure proxy (N=86), mean lexical R rises 0.017→0.235 under disclosure, and a length-nearest derangement reduces S3 recoverability by 0.136 (CI [0.118,0.158]). The same identity effect transfers to AI Idea (n=3,495) and FIRE (n=35). In a blinded two-reader pilot over 60 frozen pairwise comparisons, readers agree with the continuous higher-R side on 73.3% and 68.3% of pairs (full-sample Cohen's κ=0.83). In a compact panel, endpoint scores are associated with R (same-family r=0.75; scorer-disjoint r=0.65), and the winning model changes between S0 and S3 for 65% of multi-model items—associational, not causal. FIRE combines high identifiability (gold@1 31/35) with moderate availability, indicating that the two axes capture distinct properties. We recommend validating both axes with evidence-identity interventions to delimit when endpoint scores can support claims about scientific discovery.