Language-Grounded Drug Representations for Cold-Start Drug–Drug Interaction Prediction
Abstract
Graph-based drug–drug interaction (DDI) predictors are typically evaluated transductively, with all drugs seen during training, a regime that does not reflect the clinically relevant case of a newly approved or newly characterized drug with no interaction history. We study whether fusing a molecular-graph encoder with a frozen biomedical language-model embedding of a drug's textual description improves robustness in this cold-start setting. Across two datasets (DrugBank, multi-relational, 86 interaction types; BioSNAP ChCh-Miner, binary) and three evaluation regimes, two of them cold-start (transductive, one drug unseen (S1), both drugs unseen (S2)), fusion is never the best transductive model but has the smallest transductive-to-cold-start gap of four compared models and wins outright on every cold-start column on both datasets, e.g. DrugBank S2 AUPRC 0.376 ± 0.025 vs. 0.344 ± 0.018 (graph-only) and 0.251 ± 0.010 (text-only). A structure-only fingerprint baseline is the strongest transductive model on both datasets but degrades the most under cold start, consistent with our hypothesis that text carries complementary, topology-independent signal. We report these results with the same 3-seed protocol throughout, disclosing a threshold-calibration artifact that reverses an apparent F1 result and a literature-baseline reproduction gap we could not close, rather than omitting them.