Blind spots and strengths of LLMs in predicting chemical reaction conditions and incompatibilities
Abstract
Retrosynthesis systems typically propose synthetic routes without specifying reaction conditions, leaving a gap between a route on paper and its execution in the laboratory. Using two complementary tests, we assess how well condition-proposing systems understand condition-modulated reactivity and functional-group incompatibilities. An adversarial benchmark probes whether systems recognise incompatibility traps set around functional-group context, positional relationships, and stereochemistry; frontier LLMs outperform both the naive baseline and QUARC. A forward-prediction (round-trip) test on reactions distant from the patent-data distribution shows recovery rising consistently across four regimes: permuted, absent, QUARC-predicted, and LLM-predicted conditions, with the last giving the highest recovery for every forward-prediction model and the strongest reactivity signal overall. These results suggest that most condition predictors, including specialised, patent-trained systems, still rely more on reaction-centre memorisation than on mechanistic reasoning, and motivate a shift towards condition-falsifiable benchmarks for evaluating single-step reactions proposed by retrosynthesis models.