Multimodal Molecular Representation Learning under Incomplete Cellular Responses
Abstract
Molecular representation learning is fundamental to molecular property prediction in drug discovery. Biological responses, including cell morphology, gene perturbation, and gene expression, provide functional information complementary to molecular structure, but are often sparsely and asymmetrically available across molecules. This makes it challenging to learn molecular representations that benefit from biological context without depending on its availability. We propose Multimodal Molecular Representation Learning under Incomplete Cellular Responses (MOLIC), which organizes the consistently available 1D, 2D, and 3D molecular views into a stable multi-view structural anchor and uses observed or inferred biological signals as contextual supervision during pretraining. Joint Geometric Alignment models the joint geometry between the molecular anchor and multiple biological responses, while Cross-view Molecular Reconstruction promotes information exchange among complete structural views to provide stable structural supervision. Experiments on ChEMBL2K and Biogen3K show that MOLIC consistently improves classification and regression performance over existing multimodal molecular representation methods.