LLM’s Cannot Elucidate Structure on Out-of-Distribution MS/MS Spectra
Abstract
A frontier language model was recently reported to beat specialist tools at structural elucidation from 1D NMR spectra, with no tools of its own. We test whether this result carries over to tandem mass spectrometry (MS/MS), another common way to identify unknown chemicals in a biological sample. Our 100-spectrum benchmark from MassSpecGym is split into two groups of 50 analytes that differ in how well documented they are. Without tools the model identifies 52% of the well-documented analytes and 2% of the undocumented ones. Tools that only supply spectral annotations leave that 2% untouched. A de novo generator model lifts undocumented performance to as much as 38%, but on this split the agent adds little to what the generators achieve alone. Though an agent is effective at reranking, on the undocumented analytes it struggles to propose a correct structure of its own. This evidence suggests that LLM-based structure elucidation relies on recall where the literature is thick, and has difficulty generalizing to unfamiliar analytes without specialist tools.