What Do LLMs Use to Judge Synthesis Recipes? An Experimental Analysis on 240 Solid-State Syntheses
Abstract
Inorganic materials discovery is increasingly limited by synthesis rather than by material design. Large language models (LLMs) are being adopted to propose synthesis conditions with little experimental evidence of their capabilities and biases. Here, we carry out an experimental synthesis benchmark of 240 experiments across eight known and unknown inorganic oxides and subsequently evaluate an LLM as a zero-shot ranker over fixed banks of solid-state synthesis recipes. For each pair of recipes for the same oxide target, GPT-5.6-Sol selects the recipe expected to yield higher target-phase purity. We aggregate these pairwise preferences using a regularized Bradley–Terry model to obtain a ranked list of synthesis recipes for a given target material. By comparison with the actual purity, obtained through X-ray diffraction, we are able to measure GPT-5.6-Sol's performance as a zero-shot ranker (0.72 pairwise accuracy). By selective removal of the information exposed to the model during this ranking, we conclude the ranking is driven predominantly by thermal features, mostly independent of target chemistry. Removing all chemistry information caused little average degradation (0.68), whereas removing temperature-related information substantially reduced ranking performance (0.56). We measured this across both known and unknown targets present in the benchmark and while ranking performance was higher on known targets, the group-level difference was not statistically conclusive. In a retrospective discovery campaign, the resulting rankings provide useful 'cold start' recipe prioritization, outperforming a Gaussian process surrogate. The methodological choices highlight the need for more experimentally validated synthesis benchmarks and position LLMs as a thermally informed prior.