Towards Agentic AI for Homogeneous Catalyst Design
Abstract
This work evaluates the capacity of Large Language Models (LLMs) to reason about homogeneous catalysis in organic chemistry and applies these models to the de novo design of catalysts. We first probe what an LLM understands about a catalytic reaction before entrusting it with catalyst design. Across six mechanistically diverse catalytic systems, this probe reveals a clear performance ceiling: models retrieve established textbook chemistry from parametric memory but fail to predict reaction pathways for multi-catalyst relays, dual photoredox manifolds, and single-electron pathways without explicit literature context, and predicting selectivity remains markedly harder than predicting reactivity from SMILES representations. These observations directly motivate the design of three prospective campaigns that determine when an LLM can function as a stand-alone discovery engine versus a soft-guidance layer for a complementary machine learning algorithm. Specifically, a closed loop coupling a generative model with an LLM recovers a known catalyst class for the Morita-Baylis-Hillman reaction; an LLM-only active-learning loop optimizes ligands for a reported C-H activation; and we propose a hybrid LLM-Bayesian optimization workflow targeting a cooperative dual-ligand decarbonylative fluorination on a self-driving lab platform. These results define the boundaries of language models in catalysis and provide a framework for agentic catalyst design.