Evidence, Identity, or Position? Disentangling What Drives LLM Drug Target Prioritization
Abstract
Language models are increasingly used to rank candidate drug targets, and are judged on how well their rankings match a reference. Such scores cannot show why a model ranked as it did, because three things drive the answer: the evidence supplied, what the model already knows about the named genes, and where each candidate sits in the list. We separate all three. Across 30 diseases with five candidates each, we reveal evidence in four stages, from genetics through to safety, and present every stage twice, once with real gene names and once with them hidden, while permuting the order. Because no agreed correct ranking exists, we compare each model against itself. Across eight models and 776 controlled scenarios, disclosing an entire class of safety evidence changed the top choice no more often than shuffling the same five candidates, and this non-evidential influence grows rather than shrinks as evidence accumulates. Targets with no recorded genetic evidence are ranked 1.65 of five positions lower than targets with any, with overlapping intervals across every model tested.