Interpreting Latent Disease Neighborhoods as Phenotype Hypotheses
Abstract
Hidden states from a frozen language model can identify clusters of diseases that later Human Phenotype Ontology (HPO) releases independently mark as related. We test this on 552 rare diseases that had no phenotype annotations in 2024 but gained annotations by 2026. The method finds diseases with similar hidden-state representations, copies those diseases' existing HPO identifiers, and ranks the copied identifiers without seeing the 2026 labels. Precision at 10 is 0.347, compared with 0.277 for a baseline that combines name matching with Mondo Disease Ontology (MONDO) relationships, and 0.140 for randomly chosen neighbors. The improvement is mostly from finding the neighbors, not from later reranking. When the recovered identifiers are given to LIRICAL (Likelihood Ratio Interpretation of Clinical Abnormalities), the correct disease appears in the top 10 for 29.3% of 743 patients, versus 0% when neighbors are random. Predictions are built from the 2024 snapshot and scored on 2026 additions, following the Critical Assessment of Functional Annotation (CAFA). Each hypothesized term is an existing HPO identifier backed by the donor diseases that support it.