LP-RAG: Learning to Retrieve with Link Predictors
Abstract
Retrieval-augmented generation (RAG) strategies have empowered large language models (LLMs) through integration with external knowledge sources, leading to more accurate, up-to-date, and contextually relevant outputs. Recently, graph-based RAG methods have gained attention for leveraging relational structures that support multi-hop reasoning and retrieval. However, existing approaches remain limited in their ability to explicitly model query-aware relevance signals during retrieval. In this paper, we propose LP-RAG, a graph-based framework for document RAG that casts retrieval as an inductive link prediction problem. In particular, LP-RAG first constructs a graph encoding semantic relationships among chunks (i.e., candidates for retrieval), which is then augmented with chunk-conditioned synthetic queries that emulate potential user questions. This enables self-supervised learning of query-chunk relevance without relying on real query annotations. Retrieval is performed by predicting links between unseen queries and chunk nodes. Notably, LP-RAG is model-agnostic and can leverage a broad class of link prediction methods. To demonstrate its effectiveness, we evaluate LP-RAG across diverse settings and benchmarks. Our empirical results show that LP-RAG consistently outperforms existing methods in both retrieval quality and downstream generation performance, while maintaining competitive efficiency relative to learnable graph-based RAG approaches.