Every Measurement, Every Direction, All at Once: Multimodal Flow Matching for Molecules and Spectra
Thorben Prein ⋅ Elton Pan ⋅ Rafael Gomez-Bombarelli ⋅ Santiago Miret
Abstract
Generative modeling for science is data-scarce: each sample is expensive but arrives with several coupled measurements of the same system. Standard conditional models often use single-direction generation, discarding supervision from the rest. We argue that this asymmetry should not be a constraint during training: Under random target/conditioning masking, every measurement can act as a training target, hence one expensive sample yields multiple supervised denoising tasks. We present MOSAIC, a multimodal flow matching transformer over molecular 3D structure and UV, IR, and Raman spectra. On QM9S, MOSAIC achieves state-of-the-art performance with 88.8% Acc@1 for spectra-to-structure task, outperforming the strongest baseline by over 20% with 20$\times$ fewer integration steps. MOSAIC enables any-to-any generation from any combination of input to any combination of output modalities common in chemical sciences. We provide ablations to demonstrate the effectiveness of MOSAIC and show consistent performance scaling with increasing model and data size.
Chat is not available.
Successful Page Load