Chemical Reasoning Transfers Across Familiar and Unseen Reaction Notations
Abstract
Chemical reasoning by large language models is tested almost entirely through notations that were abundant in their training data. Whether the underlying knowledge is tied to those notations remains unknown. To separate knowledge from notation, we hold 100 named reactions fixed and rewrite each as reaction SMILES, a retrosynthetic template, and a canonical reaction signature postdating every evaluated model's knowledge cutoff. The signature's grammar is either supplied or withheld. Ten of eleven models name reactions from SMILES at 98-100 out of 100. From the unseen signature they reach 61-95 with its grammar supplied and lose up to 47 points without it, differing sharply within model families. Every model also names reactions correctly while misreading their bond changes. Applying a supplied grammar and inducing one are therefore distinct capabilities, and equal scores can hide different routes to a name. Chemical-notation benchmarks should score structural recovery alongside labels.