Multi-Document De-anonymization in Graph RAG: The MRRA Threat Model and GRAND-RAG, a Graph-Aware Hierarchical Defense
Abstract
Retrieval-augmented generation systems can disclose private facts that are not stated in any single document, but that become inferable once relations are combined across documents. We formalize this threat as the Multi-document Relational Reconstruction Attack (MRRA), in which a black-box agent iteratively reconstructs an entity-relation graph from system responses and plans its next query by reasoning over that graph rather than over recovered text. To quantify the resulting risk, we introduce DCR@q, a de-anonymization connection ratio indexed by query budget, bounded in path length and inference confidence, and reported jointly with attacker precision. On a benchmark derived from HotpotQA, MRRA attains DCR@10 = 34.7% against a Graph-RAG index, exceeding chunk-level agentic reflection by 11.3 points (p < 0.001); because the two attackers differ only in the object of reflection, this gap isolates the contribution of graph-awareness. We further propose GRAND-RAG, an inference-time defense that gates generated output on the marginal increase in DCR. It reduces DCR@10 to 0.0% while retaining 99% of undefended supporting-fact recall at a cost of at most 11 ms per query, provided that the gate is applied incrementally within each response; applied to the response as a whole, it leaves DCR@10 at 22.7%. Our central finding concerns Korean. Agglutinative morphology and the coexistence of NFC and NFD encodings fragment 36% of entity mentions in real text. We prove that such fragmentation biases the defender's risk estimate downward (Proposition 1) and confirm this without exception across 241 targets, where 79% of completed private paths are rendered invisible to the defender's own metric; we further show that the same fragmentation disables the relational gate itself (Corollary 1). Finally, we report negative results: the relational retrieval layers reduce false refusals but do not improve privacy, the session-level risk budget yields no measurable benefit, and neither our defense nor corpus rewriting protects private pairs that the defender has not enumerated in advance.