Tandem Mass Spectra Are Not Text: The Dependence of LLM Structure Elucidation on Domain Tools
Abstract
A frontier model was recently reported to solve 1D NMR structure elucidation unaided, returning published structures from a peak list and a molecular formula. We ask whether this capability transfers to tandem mass spectrometry (MS/MS), where fragmentation is stochastic, instrument-dependent, and has no textual analogue in pretraining. To measure elucidation of novel analytes, we separate spectral inference from literature recall and derive a 100-spectrum benchmark from MassSpecGym whose two arms are matched on acquisition conditions but differ in literature footprint. We hold the harness, model, and output contract fixed, and vary only the model’s access to domain tools: subformula annotation, forward spectrum prediction, and de novo generation. We find that unaided LLM Top-1 accuracy is 52% on heavily documented analytes and 2% on undocumented ones. Only de novo generator tools raise performance on undocumented analytes, to as much as 38% Top-1 accuracy. This gain is bounded by the tool, since on undocumented analytes the agent almost never returns a structure its generator did not propose. Thus, performance tracks the domain tools rather than the reading of the spectrum.