When One Diagnosis Is Not Enough: Explanatory Completeness in Rare-Disease Diagnosis
Abstract
Agentic diagnostic systems are increasingly evaluated by recovery of a single reference disease, even though some patients require a set of molecular diagnoses to explain their clinical presentation. We evaluated whether phenotype-driven diagnostic systems recover these multiple diagnoses using synthetic dual-diagnosis profiles that we constructed and multilocus cases curated from the literature. Both diagnoses appeared in the Top-5 for only 4.9-9.7% of synthetic and 0-4.2% of literature-derived dual-diagnosis cases. In the agentic diagnostic system (DeepRare), many reference diagnoses were never identified during retrieval. Even when both reference diagnoses were supplied as candidates, a shared downstream decision prompt selected only one diagnosis in 62.2% of synthetic cases. Explanatory-completeness checking reduced single-diagnosis selection and improved complete retrieval in the synthetic LLM-only experiment. However, this retrieval improvement did not extend to literature-derived cases, and predictions of additional diagnoses increased in controls labeled with a single diagnosis. These findings highlight limits of current phenotype-driven diagnostic workflows and motivate approaches that explicitly model diagnostic completeness. We release synthetic and literature-derived multilocus benchmarks to support the development and evaluation of the next generation of LLM-based and agentic methods that recognize when one diagnosis is insufficient.