The Retrieval-Robustness Paradox: Data Poisoning in Retrieval-Dependent Domains
Abstract
Retrieval-augmented generation (RAG) has been critical in enabling large language models (LLMs) to provide up-to-date and domain-specific answers by retrieving external documents. However, the reliance on external documents introduces a new security risk: an attacker can inject malicious documents to manipulate the output of a RAG system. In response, many defenses were proposed to mitigate attacks. In this work, we perform a critical evaluation for state-of-the-art defenses. Specifically, existing defense evaluations were performed on retrieval-redundant (RR) tasks, where external retrieval is unnecessary because LLMs already know the answers, leaving a gap between measured robustness and real-world application scenarios of RAG, where correct answers are only obtainable through the retrieved documents. To bridge the gap, we present a systematic evaluation framework that distinguishes RR tasks from retrieval-dependent (RD) tasks. Across six QA tasks, we reveal The Retrieval-Robustness Paradox: defenses recover 26.45 percent less in the RD tasks where RAG matters most. We observe an Internal Fallback Effect, where perceived retrieval robustness in RR tasks is actually the result of the model already knowing the answer. Our findings highlight the need for developing robust RAG systems in RD tasks. More broadly, our work showcases how evaluations on RR tasks can create an illusion of robustness that fails to generalize to the RD tasks where RAG matters most.