Choosing Experiments, Replication, and Recalibration under Measurement Drift
Abstract
Scientific agents must decide how to spend a fixed measurement budget when the assay drifts: measure a new candidate, replicate, query a reference, or recalibrate. We contribute an auditable, pre-declared evaluation harness and protocol for scripted measurement-management policies, with typed tools, credit enforcement, a separate evaluator, and hashed ledgers; it is not evidence about LLM agents. Naively, earlier studies looked puzzling: on 96 analytic surfaces a two-action knowledge-gradient lookahead differed from a tuned adaptive controller by −0.00033 regret (95% range −0.00574 to 0.00449) and never queried a reference, and in a metabolomics archive four controllers made identical choices for 28 of 28 features. The pre-declared control (E4) uses 9,600 fresh surfaces with one unannounced offset jump per episode, a positive control, static, random and oracle comparators, and action-removal arms. Managing measurements roughly halved regret versus candidate-only selection (0.01022 versus 0.02038; both Holm-tested contrasts detected), which validates reset utility in this simulator, not reference utility or physical recalibration. Disabling the reference trigger had no detectable effect, whereas disabling reset substantially worsened regret. Lookahead was practically equivalent to a budget-aware threshold controller (difference 0.00103; 90% interval inside the pre-declared margin δ = 0.005, derived from saved-data effects; pre-declared minimum detectable effect 0.0033) at a per-episode CPU-time ratio of 143. Limits: no evaluated controller obtained reference value; lookahead, even with a jump-aware drift model, never queried references, so that reference ablation is identically zero; lookahead was worse in a secondary abrupt-drift panel (descriptive); and no language model, robot, or laboratory experiment is tested.