Benchmarking Bayesian optimization strategies across compositional landscapes
Abstract
Bayesian optimization (BO) is increasingly used in chemistry and materials science to accelerate experimental discovery. Gaussian process (GP) surrogates and expected improvement are the most frequently reported choices in experimental BO workflows, but their suitability depends on the generally unknown optimization landscape. Previous benchmarks have focused on surrogate–acquisition pairings while overlooking initial sampling and batch size; some also use datasets partially collected under BO guidance, complicating retrospective comparisons with random search. Using twelve high-throughput metal oxide datasets spanning ternary and quaternary composition spaces, we benchmark practical BO design choices, including surrogate model, initialization strategy, and batch size. We find that random forest (RF) surrogates generally outperform GPs across the primary benchmark and are particularly robust when only a few initial observations are available, whereas GP performance improves substantially with additional initial data and configuration tuning. In the ternary benchmarks, using about ten initial experiments offered a practical balance between model stability and sample efficiency, while batches of five reduced the number of closed-loop cycles with a modest increase in the total number of experiments. Overall, our results show that BO performance varies across compositional landscapes and provide practical guidance for adapting BO strategies to the available data, optimization objective, and experimental throughput.