Adding a Biomedical Knowledge Graph Can Make an LLM Worse at Drug Repurposing
Abstract
Graph-based retrieval augmentation (GraphRAG) is widely assumed to make large language models (LLMs) more reliable on domain tasks by anchoring predictions in curated facts. We test this assumption on drug repurposing for Parkinson's disease with three controls that prior evaluations have not combined: an unaugmented baseline, corrupted-graph controls, and a strict temporal holdout (T = 2016-07-01) over a drug-target-disease knowledge graph (KG) built from ChEMBL and Open Targets. Scoring 90 drugs with Qwen3-8B, the model alone reaches AUROC 0.807; the real KG lowers it to 0.562 (paired bootstrap A = -0.243, 95% CI [-0.466, -0.021]), and shuffled and target-permuted graphs lower it further, mono-tonically. Citation grounding stays at 73-74% under every graph, so the model reads and defers to what it retrieves, and is misled by it. We trace the failure to a representational property: association encodes which proteins are relevant to a disease, not whether a drug should raise or lower their activity. Association alone ranks parkinsonism-inducing D2 antagonists above every drug that entered a PD trial, a denser graph performs worse, and the same mismatch governs a graph neural network on breast cancer and glioblastoma. Grounding and accuracy are orthogonal; KGs for repurposing must encode direction, and evaluations must include unaugmented and temporally clean controls.