Exposing Private Corpus Leakage in Multimodal RAG
Abstract
Multimodal Retrieval-Augmented Generation (RAG) helps mitigate hallucinations in Vision-Language Models (VLMs) by grounding generations in external knowledge bases. However, this externalized memory also introduces privacy risks, as external queries may reveal signals about sensitive visual records in the retrieval corpus, such as medical images, scanned contracts, and proprietary business documents. This exposes a retrieval-corpus privacy risk: private records may be detectable through interactions with multimodal RAG systems. To study this risk, we propose the Semantic Degradation Attack (SDA), a two-query method for exposing private corpus leakage in multimodal RAG by testing how strongly generated responses depend on retrieved evidence. SDA constructs transferable retrieval-disrupting perturbations using local surrogate visual encoders. By measuring the drop in the semantic similarity of the VLM's generated descriptions before and after perturbation, SDA can distinguish whether a target image-caption record exists in the private retrieval corpus. Extensive experiments demonstrate that SDA consistently detects private corpus leakage more effectively than existing baselines on two image-caption datasets across five popular VLMs.