Persistent Homology with Vineyards for Robust NMR Representation
Abstract
Automated molecular structure elucidation from experimental Nuclear Magnetic Resonance (NMR) spectra is fundamentally limited by noise, baseline distortions, and overlapping signals. Existing neural encoders either read the spectrum as a dense grid of intensities, directly entangling spectral artifacts with the chemical signal, or reduce them to picked peak lists, placing the entire representation downstream of a single thresholding decision. To address these limitations, we introduce VINEA, a deep learning framework that encodes 1D spectral topology using persistent homology. Rather than relying on rigid spatial grids or heuristic peak-picking thresholds, VINEA extracts topological features across a Gaussian scale-space vineyard, using a fibered barcode to distinguish persistent structural multiplets from transient noise, and decorates them with geometric properties including coupling-sensitive separations in Hz. The resulting unordered token set is processed by a Set Transformer in which the attention bias depends only on pairwise chemical-shift differences, making it invariant to global referencing drift. At a fraction of baseline parameter counts, VINEA strictly outperforms existing methods, achieving state-of-the-art accuracy on clean spectra, heavily degraded signals, and transfer to real experimental spectra.