Attention-guided interpretable link prediction for heterogeneous biomedical knowledge graphs
Abstract
Link prediction in biomedical knowledge graphs (KGs) is crucial for discovering novel disease-gene associations, yet noise in the graph and the lack of proper interpretability tools limit its utility. We introduce EchoLink, an interpretable heterogeneous graph transformer (HGT) with sparse top-K attention designed to address uncertainty in existing edges and provides both subgraph- and path-level explanations for its predictions with minimal computational cost. On the Open Targets biomedical knowledge graph, EchoLink outperforms every baseline tested, exceeding the strongest by 5.6 points in Recall@100, recovering 3.48×as many targets among sparsely connected genes, and, when frozen on the 2021 release, successfully predicting 28.6% of future clinically connected targets within its top 100. Our attention-based interpretations are comparable to expensive perturbation methods on motif-detection benchmarks while being two orders of magnitude faster on average. Applied to the Open Targets KG, our explanation method successfully extracts disease-relevant mechanisms that are validated by scientific literature for 16 of the top 50 novel predictions. EchoLink thus enables scalable and interpretable prediction that can accelerate novel therapeutic target discovery.