Bayesian Low-Rank Posteriors for Scalable Membership Inference
Abstract
Membership inference attacks (MIAs) aim to determine whether a sample was used during the training of a target model. The most effective MIAs rely on reference (or shadow) models to estimate target-conditional score distributions, but training such ensembles is computationally prohibitive for modern large-scale Vision--Language Models (VLMs). In this work, we propose a scalable alternative that replaces explicit reference-model training with a Bayesian approximation over low-rank adaptation parameters. Leveraging parameter-efficient fine-tuning, we model uncertainty only within the LoRA subspace and construct a posterior from a single training run, from which we sample a diverse set of virtual reference models. We instantiate this approach using stochastic weight averaging Gaussian (SWAG), enabling efficient approximation of target-conditional score distributions at a fraction of the cost of conventional shadow-model ensembles. We evaluate the resulting attack on three downstream-adapted VLMs across two privacy-sensitive visual question answering tasks, namely Document VQA and medical VQA. To isolate the effect of pretraining knowledge, we additionally introduce a controlled synthetic dataset. Our approach outperforms standard reference-based attacks while requiring only a single trained reference model, demonstrating that accurate and scalable membership inference is feasible even for large VLMs.