Did the Model Forecast It, or Already Know It? Evaluating Large Language Models for Novel Drug-Target Prioritisation with a Temporal Knowledge-Graph Oracle
Abstract
AI systems increasingly prioritise which drug targets enter costly clinical validation, but it remains unclear whether they anticipate genuinely novel target--disease links or restate ones already in their training data. Retrospective benchmarks cannot tell these apart: they score models against associations that predate the evaluation and retrieve evidence with no decision-time bound. We evaluate six open language models on a temporal biomedical knowledge graph rewound to each decision point, ranking the targets a disease will newly enter trials for before the supporting evidence exists. Parametric knowledge alone forecasts above chance, and that floor, not zero, is what retrieval must beat. Decomposing what retrieval adds on top, honest decision-time evidence gives only a small lift, significant for three of six models and weakening as the pair becomes more recent. The newest graph appears far stronger, but that gain is almost entirely post-cutoff \emph{clinical precedence}: deleting only those trial edges collapses it, fresh non-label biology adds nothing reliable, and only the leaked precedence resists temporal decay. Current models thus recall established target--disease precedent but do not forecast novel ones. Target prioritisation should be evaluated temporally, against the evidence available at each decision point.