Target Alignment in Sequential Molecular Model Approval
Abstract
Scientific workflows often let selected molecular observations decide whether a predictive model replaces its incumbent. A sequential test can be valid for the selection distribution yet fail to protect the broader target the workflow names. We study this in an offline molecular update-approval benchmark for future scientific agents, run by a scripted controller, not an LLM. Across three datasets, 90 of 1,980 nonimproving directed model comparisons improve under a disagreement-based evaluation law (23 beyond a 0.005 loss margin). Across thirty six-round campaigns per policy, the unweighted selection gate makes 78 promotions, ten worsening the declared uniform target (one beyond the margin); uniform auditing (56) and importance weighting (25) make none, and an exploratory prediction-specific bound raises importance-weighted promotions to 59. Seven of the ten target-worsening promotions improve external loss. Post hoc, gates read on average at least 95.9% of the distinct target molecules. A pre-declared label-costed study (80 split units) hides archived target labels until bought and shares a budget of 25% of the target across rounds. The uniform audit stops on budget in 476 of 480 rounds; spending the same labels on training gives lower final target risk (audit minus training +0.0424 clipped-loss units, 95% CI [+0.0366, +0.0484]) and external risk (+0.0456), with no unit favouring the audit. The audit’s zero observed target and external harms coincide with 4 promotions in 480 rounds, as a pre-freeze pilot had shown: in this fixed six-round, fixed-α/6 schedule it was label-starved, and cheaper labels, budget-scaled schedules and larger per-update gains remain untested.