Predicting Responses to Unseen Drugs Using Knowledge Graph-Enhanced Biomedical Text Embeddings from Large Language Models
Abstract
Predicting responses to previously unseen drugs remains a major challenge in drug response prediction (DRP), because response measurements for new compounds are unavailable during model development. Meanwhile, rich biomedical knowledge describing drug mechanisms and cellular contexts is available, but incorporating such knowledge directly into prediction models can make inference dependent on heterogeneous and potentially incomplete external information. We instead formulate biomedical knowledge as privileged semantic supervision available during training but not required at inference. For each drug–cell pair, we retrieve mechanistic relationships connecting drugs, targets, pathways, and cell-specific molecular contexts, augment them with response- and vulnerability-derived similarity context, and convert the resulting knowledge into biomedical text. A pretrained biomedical language model encodes this pair-specific knowledge, which is aligned through contrastive learning with an experimentally derived drug–cell representation used for response prediction. At inference, the knowledge branch is removed and prediction relies solely on the aligned experimental representation. Under 10-fold drug-blind evaluation on GDSC2, our framework improved PCC from 0.657 to 0.684 and reduced RMSE from 1.850 to 1.750 over the experimental backbone. Ablation and representation analyses further support the complementary role of biomedical knowledge and cross-modal alignment. These results suggest that biomedical knowledge can improve unseen-drug generalization by serving as training-time semantic supervision rather than an inference-time dependency.