When Scientific Simulators Disagree: A Falsifiable Audit for Surrogate Verification in AI for Science
Abstract
AI-for-science systems increasingly use simulators as scalable verifiers, but a simulator is a model rather than ground truth. We advance a contestable position: a high score from one imperfect scientific surrogate is insufficient to authorize an AI-generated conclusion unless the evaluation exposes verifier disagreement, endogenous scoring, objective-specific counterexamples, and claim boundaries. We instantiate the argument with two transparent but simplified planetary-feedback bridges: a World3-style system-dynamics implementation with 630 policy episodes and a DICE-style climate-module implementation with 2,525 episodes. Removing the direct author-defined policy-prior term preserves the reported aggregate winner in both bridges. A frozen exploratory grid that scales every score component by 0.5, 1.0, or 1.5 also preserves that winner at all 6,561 World3-style and 2,187 DICE-style grid points. Yet objective-specific verification disagrees: specialized policies outperform on overshoot, resource preservation, stability, peak temperature, regret, or smoothness. Thus aggregate robustness does not imply objective-invariant scientific validity. We propose a minimal surrogate-verification report and explicit falsifiers. This is a position paper supported by post hoc exploratory audits, not a new learned model, official World3/DICE evaluation, real climate validation, or policy recommendation.