PrivCap: Benchmarking Cross-Modal Interference in Vision-Language Privacy Detection
Abstract
Privacy in social-media posts requires assessing images and captions together. However, existing benchmarks largely evaluate image and text privacy separately, and recent multimodal datasets often depend on generated images or fully synthetic user profiles. We introduce PRIVCAP, a matched benchmark pairing 10,391 real images from VISPR, Privacy Alert, and DIPA2 with one private and one safe natural social-media caption each (20,782 posts) under a unified 14-category privacy taxonomy. Our controlled design enables us to isolate the effect of text on visual privacy detections by evaluating the same image under both safe and private captions. We evaluate four open-weight vision-language models (3B-11B) across detection, recognition, and modality attribution. We find that text captions interfere with the models: 10 of 12 settings show models become worse at detecting visual privacy leaks in the image when a private caption is added (drops up to 29.3 F1 points). And while models can often tell that a post is private, attribution of the leak to the correct modality exceeds 35% in only two settings, both on VISPR.