CaLMBO: Candidate Selection with Language Models for Bayesian Optimization in Closed-Loop Scientific Experiment Design
Emma Pajak ⋅ Maximilian Bloor ⋅ Laura Marie Helleckes ⋅ Antonio del Rio Chanona
Abstract
Autonomous laboratories must choose informative experiments from expensive design spaces under budgets of tens to hundreds of trials. Bayesian optimization (BO) is a well-established approach for sample-efficient optimization in this setting, but typically operates on numerical decision variables without access to their scientific meaning. Large language models (LLMs) can condition on rich natural-language descriptions, providing a natural interface for incorporating scientific context within optimization. We present CaLMBO (CAndidate selection with Language Models for Bayesian Optimization), which places an LLM in an expert-in-the-loop role within BO: a batch acquisition function generates a candidate set, and the model selects among these candidates using scientific context. Across four scientific optimization problems (8$-$18D, 30 runs each), CaLMBO attains the highest mean normalized AUC across the compared methods, with the highest per-problem mean on the two highest-dimensional problems and overlapping 95\% confidence intervals with the best method on the other two. Matched ablations give the ordering CaLMBO $ > $ CaLMBO-LHS $>$ BO-UCB on all four problems, showing that both in-loop selection and LLM-guided initialization contribute. Removing scientific context from the LLM degrades performance on all four problems.
Chat is not available.
Successful Page Load