Per-Element Spectral Conditioning Addresses a Bottleneck in NMR-Conditioned Graph Diffusion
Abstract
Automated identification of chemical compounds requires generative models that turn spectroscopic measurements into exact molecular topologies. Diffusion models for Nuclear Magnetic Resonance (NMR) structure elucidation compress the whole spectrum into a single pooled vector and apply it identically to every atom and every bond, through feature-wise linear modulation (FiLM) or adaptive layer normalisation (adaLN). That interface is treated as settled, and we find it is a binding constraint in our setting: replacing it recovers structures that the encoder changes we tested did not. We introduce CaNMR (cross-attention NMR conditioning), which replaces both FiLM modulations in the DiGress graph-diffusion backbone with spectral cross-attention, so that each atom and each candidate bond queries the spectral tokens for itself at every layer. For a single conditioning step on the edge stream, the functions CaNMR can express strictly contain those FiLM can. On the SpectraBase benchmark of simulated spectra, with the molecular formula fixing the atoms so that the task is bond assignment, top-1 exact structure recovery rises from 43.26\% to 52.89\%. That is a gain of 9.63 percentage points for 2.3\,M additional parameters on our 16.6\,M reproduction of the NMR-DiGress baseline. Adding an auxiliary spectral reconstruction loss to a trained CaNMR model reaches 54.87\%. Added pooling capacity does not explain the gain: pooling with three learnable queries rather than one, into the same conditioning width, recovers only 1.47 of the 9.63 points. The gain is not free: chemical validity falls from 94.2\% to 90.4\%. Fixing the conditioning interface is a one-off, 14\% parameter change, and it gives a route to better structure recovery with minimal scaling.