MolDisBench: Grounding and Calibration\\ in LLM Agents for Molecule--Disease Reasoning
Sibasankar Panigrahy ⋅ Ankit Patidar
Abstract
Large language model agents are increasingly deployed for drug repurposing and toxicological triage, tasks that reduce to classifying molecule--disease relations. Existing benchmarks score systems predominantly on curated positive relations, confounding evidence retrieval with hazard-class pattern matching. We introduce MolDisBench, a 47-item benchmark of molecule--disease links verified against primary regulatory and trial sources and stratified by evidential status into positive, near-miss, falsified and contested relations. Near-miss items pair the molecule of one positive item with the documented outcome of another matched on hazard designation, holding class-level plausibility largely constant so that molecule-specific evidence is the primary discriminative signal. We evaluate ten models from four providers under three evidence conditions: parametric knowledge alone, a supplied reference passage, and agentic search. Supplied and self-retrieved evidence yield distinct performance profiles. A fixed passage yields 89--96\% accuracy for seven of eight models regardless of their parametric-knowledge baseline, whereas agentic search improves the strongest model by 6 points and degrades the weakest by 7, below its own parametric score. Stratum decomposition shows that the positive and falsified strata are non-discriminative: every model scores $12/14$ and $10/10$ on them and fails on the same two positive items. Discriminative variance is confined to the near-miss (29-93\%) and contested (33-67\%) strata, across which inter-model agreement falls from $\kappa{=}0.79$ to $0.54$. We further formalize abstention-based label auditing: the chance of spuriously corroborating a sound label increases with panel size under permissive thresholds, and correlated abstentions ($r{=}0.62$) reduce nine models to $k_{\mathrm{eff}}{=}1.51$ effective votes.
Chat is not available.
Successful Page Load