Surrogate Verifiers Shape Decisions in Closed-Loop Experimental Design
Abstract
High throughput, automated screening has revolutionized biological experimentation and prompted predictions of AI-enabled "self-driving" labs. Deployed autonomous lab systems typically use a learned surrogate to predict new experimental candidates and rank which conditions deserve expensive follow-up. Determining whether a computational surrogate verifier can substitute for an experimental assay’s ground truth requires an understanding of model accuracy, uncertainty, and extrapolation. In this work, we present surrogate verifiers trained on N=7,688 Lacticaseibacillus growth curves exploring a 10-dimensional media component experimental design space. We train surrogate models across six distinct families including polynomial, Gaussian-process, tree-ensemble, neural, and Monod-kinetic forms to predict maximum optical density (OD600). Surrogates ranked by the most commonly reported accuracy metrics (held-out R²; MAE; RMSE) disagree sharply on which experiments to test next and the best achievable yield. We evaluate models in a Bayesian optimization loop and find that the model used as the active learning oracle matters more than the acquisition function or batch size. We apply conformal calibration to build a model-agnostic assessment criterion of uncertainty and model coverage. We end by evaluating model accuracy and calibration under transfer to an unseen related bacterial strain, and confirm on a prospective validation plate that the transferred verifier’s intervals fail in the direction and region our retrospective analysis predicts. This work explores the surrogate verifier within the loop as a previously underexplored axis of closed-loop design and points towards best calibration practices.