Multimodal Fusion for Accurate Molecular Property Prediction
Abstract
Molecular property prediction is challenging because the molecular characteristics relevant to prediction vary across tasks. Multimodal fusion has been widely reported to close this gap, yet prediction accuracy often plateaus, and the effect of fusion itself has rarely been studied systematically. To evaluate the effect of modality fusion independently of the backbone encoder, we analyze all 15 non-empty subsets formed from four modalities---3D conformer geometry, molecular graph, SMILES (Simplified Molecular Input Line Entry System) text, and molecular image---across six regression tasks. With a scratch-trained SchNet backbone, adding modalities improves the mean prediction accuracy by (+0.040) in (R^2) relative to the backbone alone. Replacing SchNet with the pretrained Uni-Mol2 backbone and repeating the entire modality lattice eliminates the average gain and produces the largest performance decreases on the two tasks for which Uni-Mol2 alone performs best. Through modality-effect analysis, we find that the performance gains from multimodal fusion are backbone-dependent and consistent with the degree of representational overlap. These results establish that the primary value of multimodal fusion lies in mitigating worst-case performance and improving the reliability of predictions across properties with varying representational requirements.