Inverse Bayesian Optimization Recovers Time-Resolved Search Strategies\\from Human and LLM-driven Experimental Campaigns
Abstract
LLMs are increasingly advocated to be placed in charge of experiment selection in scientific discovery loops, yet their choices are almost always judged by its performance. Performance metrics leave the strategy behind a campaign unobserved, so it stays unclear which decisions drove a campaign to succeed or fail. Inverse Bayesian optimization (iBO) directly recovers explore-exploit optimization strategies rather than merely scoring strategies by outcome, but existing formulations target smaller continuous tasks and are impractical for high-dimensional biochemical and material experiment optimization. We rebuild the estimator for combinatorial categorical spaces with Gaussian Process surrogates and use it to compare LLM and human search strategies on reaction-optimization campaigns. Human experts split their weight between varying one factor at a time and local search around the current state, while LLMs are greedy on the surrogate mean. Neither places meaningful weight on uncertainty-driven, explorative policies. LLMs keep their strategy when warm-started on human trajectories, indicating the behavior is inherent to the model. Because iBO fits per optimization step, policies can be monitored online and interventions timed to what the agent is currently doing. Our prompt intervention experiments show that what the agent is shown, not what it is instructed to do, is what moves its strategy.