RepurposingBench: Evaluating Scientific Hypothesis Generation Beyond Recall in Drug Repurposing
Mark Hsia ⋅ Christian Gensbigler
Abstract
Drug repurposing requires searching a combinatorial space of existing drugs and diseases for a small set of hypotheses worth testing experimentally. Large language models are increasingly capable of searching biomedical literature and reasoning over disease biology, but evaluating whether they can generate genuinely new therapeutic hypotheses is difficult because known drug-disease associations may already occur in their training data. We introduce RepurposingBench, a benchmark designed to separate retrieval of known associations from reconstruction of hypotheses from prior scientific evidence. The benchmark contains a rediscovery tier of 143 drug-disease pairs that frontier models fail to recover under targeted knowledge probes and a stricter prospective tier of 10 associations first publicly disclosed after a fixed date cutoff. Models only receive information available before each association was disclosed and rank candidate drugs while explaining their reasoning. Across seven frontier models, weighted recall on the rediscovery tier is 35-41\%, whereas no model recovers a prospective target within its top 40 candidates and only two of ten are recovered at $k=50$. Prospective rationale scores are dominated by disease biology, with almost no credit from candidate-specific pharmacology, mechanism, or clinical reasoning. The results reveal a gap between possessing biomedical knowledge and composing that knowledge into new therapeutic hypotheses.
Chat is not available.
Successful Page Load