Benchmarking Membership Privacy Risks in Preference-Based LLM Post-Training
Abstract
Modern language models are commonly adapted after pretraining to follow instructions, align with user preferences, and improve deployment behavior. Such post-training often relies on preference data from users, annotators, or model interactions, which may contain sensitive prompts, private responses, proprietary tasks, or confidential judgments. Understanding the privacy implications of this data is therefore essential. Membership inference attacks (MIAs), which test whether a record was used for training, are the de facto standard for empirical privacy auditing in machine learning. However, preference-based post-training changes the unit of membership: a training record contains a prompt, a preferred response, and a dispreferred response, rather than a single input-output sequence. Audits that ignore this structure can under-report leakage. We therefore introduce a benchmark protocol that adapts strong reference-model MIAs to complete preference records. Our protocol asks whether membership is exposed by either response alone, by the model’s preference between them, or by the two responses jointly. We find that auditing the response pair can reveal membership evidence missed when the record is reduced to a single score. Across three preference datasets, two model families, and seven post-training objectives spanning three post-training families, measured leakage depends strongly on both the audited record statistic and the training objective, with imbalanced risk between preferred and dispreferred responses. We further evaluate the impact of parameter-efficient fine-tuning (PEFT) and differential privacy (DP), finding widely varying privacy-utility trade-offs. Overall, preference-based post-training leaks in ways that standard single-response audits can miss, motivating formal privacy methods that protect complete preference records while preserving post-training utility.