Beyond Raw Context Transfer: Representation-based Federated Retrieval-Augmented Generation
Abstract
Retrieval-augmented generation (RAG) improves the factuality of large language models (LLMs) and vision-language models (VLMs) by grounding generation in external knowledge. However, most existing RAG frameworks assume a centralized retrieval document pool, which is often impractical in sensitive domains such as healthcare, where data are inherently distributed and subject to strict privacy constraints. Recent efforts on decentralized RAG primarily follow prompt-based paradigms that exchange raw, human-readable retrieved content, leading to substantial communication overhead and potential privacy risks. To address these limitations, we propose Representation-based Federated RAG (FedRepRAG), a decentralized RAG framework that avoids raw-data exchange during retrieval. FedRepRAG retrieves information from private local document stores and exchanges only compressed latent representations. To integrate the retrieved knowledge, we introduce a lightweight collaboratively trained projector that maps retrieval embeddings into the feature space of a frozen LLM/VLM backbone. This design makes the exchanged information interpretable only within the federated system, thereby enhancing privacy while substantially reducing communication cost. Experiments on decentralized visual question answering (VQA) and question answering (QA) benchmarks show that FedRepRAG consistently outperforms direct inference and local retrieval baselines, while substantially reducing retrieval tokens, computational cost, and inference-time communication overhead compared with raw-context transfer. Overall, FedRepRAG establishes an effective, efficient, and privacy-preserving framework for federated RAG.