PRISM: Prior Relational Information for Self-supervised Modeling to Enhance Solubility OOD Generalization
Abstract
Solubility depends on intermolecular interactions between solutes and solvents and is central to drug discovery, chemical and materials process engineering. Its prediction requires out-of-distribution (OOD) generalization to unseen solutes, solvents, and solvent mixtures, where supervised methods suffer from data scarcity, and existing self-supervised pretraining methods do not explicitly learn chemically meaningful relations. Learning chemically meaningful relational information without labels can alleviate data scarcity and improve OOD generalizations, but remains challenging. We introduce Prior Relational Information for Self-supervised Modeling (PRISM) and apply it to solubility OOD generalizations. PRISM is a label-free pretraining framework that converts expert-defined chemical priors into relational teachers. For solubility prediction, it trains four encoders to preserve the population-level relational information from priors capturing hydrophobic, electrostatic, polarizability, and substructural effects, using Kullback-Leibler (KL) divergence to match pairwise distance-based distributions during pretraining. For scaffold, solute, and solvent OOD generalizations, PRISM matches or surpasses supervised and self-supervised baselines. For mixed-solvent OOD generalizations, freezing the four PRISM encoders and training a light fusion-and-composition head outperforms the strongest supervised baselines and greatly reduces prediction variations. Ablations and control studies confirm that gains stem from chemically meaningful relational structures baked into the encoders by PRISM. We expect this strategy of turning relational priors into self-supervision signals to extend beyond solubility prediction to domains where chemically or biologically meaningful relational priors are available.