LLM Agents for Hypotheses Generation and Experimental Design in Heterogeneous Catalysis
Abstract
Heterogeneous catalyst discovery often operates in a low-data regime, where campaign data leave potentially important relationships unresolved. Bayesian optimization efficiently identifies high-performing catalysts, but its adaptive trajectory does not necessarily resolve why they perform well: support, composition, loading and promoters can change together across successive experiments. We ask whether LLM agents can convert questions left unresolved by a completed optimization campaign into experimental proposals whose evidence and design commitments can be checked before synthesis. We introduce a traceable representation that links each question to campaign evidence, competing hypotheses, a multi-well experimental plate and outcome-dependent predictions. We apply it to a completed campaign of 143 catalysts using 10 independent runs of each of 3 agentic coding systems. The agents converted confounded trends, missing comparisons and non-estimable interactions into candidate plates. Repeated runs characterized both recurring questions and system-dependent scientific priorities. The representation enabled separate audits of campaign grounding, chemical priors, experimental design, and predictions. These audits supported most invoked chemical priors and localized remaining problems to specific evidence and design components. We plan to submit selected plates for prospective testing on the same automated platform that produced the campaign.