The Translation Gap in Agentic Structure Elucidation: Accuracy Tracks Documentation, Not Spectra
Abstract
In drug discovery the compounds that often matter most sit in no spectral library and no structure database, forcing de novo structural elucidation. A frontier model was recently reported to solve 1D NMR elucidation without tools, from a peak list and a molecular formula. We ask whether this finding reaches tandem mass spectrometry (MS/MS), where fragmentation is stochastic, instrument-dependent, and has no textual analogue in pretraining. To separate spectral inference from literature recall, we derive a 100-spectrum benchmark from MassSpecGym whose arms are matched on acquisition conditions but differ in literature footprint. One arm contains 50 highly documented structures, many of them approved drugs or clinical-stage candidates; the other contains 50 with no literature record and the same drug-likeness. Holding the harness, model, and output contract fixed, we vary only access to domain tools. Without them, Top-1 accuracy is 2% on the undocumented arm against 52% on the documented one. Only de novo generators lift that 2%, to at most 38%, and the agent almost never returns a structure its generator did not propose. Performance tracks the tools and the documentation rather than the reading of the spectrum.