From Complementarity to Actionability: Observability and Cost-Aware Routing in Drug Repurposing
Abstract
Drug-repurposing models are compared by how well they score on average, which hides whether methods built on different assumptions succeed on different diseases, and hides whether any such difference could be acted on at query time. We study both questions in a disease-cold-start benchmark built from PrimeKG, where four heterogeneous experts rank the same 7,957 drugs for 738 held-out diseases against 8,368 CLEAN indication pairs under explicit therapeutic-label masking. The strongest single expert reaches disease-macro Recall@50 of 0.629211; a disease-wise oracle reaches 0.681351, an opportunity of 0.052141 (paired disease-level bootstrap 95% CI, 0.040873–0.064247). That opportunity is positive across seven ranking metrics, two candidate-universe restrictions, and all 24 masking-protocol-by-fold comparisons, though which expert wins a given disease is materially protocol-sensitive. We then ask whether it can be acted on. Low-cost static graph descriptors do not improve on the strongest expert. A serving-cost audit nevertheless shows that substantial oracle-constrained headroom remains within modest budgets: the cost-constrained oracle reaches 0.679905 within a 1.25× budget. We then test a bounded diagnostic computed after the default expert has run. Six score-state summaries carry partial relative-performance signal (out-of-fold Spearman 0.402908 for E1), yet a prespecified nested policy makes no escalation in any of five outer folds and reproduces the default on all 570 development diseases. Every evaluated S0 threshold policy satisfies both latency budgets. Complementarity, observability, decision value, and resource feasibility are distinct conditions; satisfying the first does not guarantee the rest.