Measuring Privacy under Approximate Memorization in LLM Training
Abstract
Membership inference attacks (MIAs) are the standard tool for auditing whether personal data was used to train a language model, but existing entity-level meth- ods assume the auditor knows the exact sentence in which the data appeared. In practice, a person who suspects their data was used knows only the identifiers themselves—their name, address or phone number—not the surrounding text. We study membership inference under this approximate memorization threat model, where the attacker reconstructs plausible sentences from partial PII instead of querying the verbatim record. Building on the EL-MIA benchmark, we build three paraphrase datasets from AI4Privacy (semantic-preserving, semantic-altering, PII-reordered) and evaluate loss, Min-K% probability and perplexity-ratio attacks against a Pythia-2.8B model trained on the original sentences. Membership stays detectable without the training sentence: semantic-preserving paraphrases reach 98.8% AUC, and paraphrases that discard the original meaning or attribute order remain well above chance (73–75% AUC). Attack strength depends on how para- phrase scores are aggregated, grows with the number of sensitive attributes per sentence, and is highest for high-entropy identifiers such as IBANs and cryptocur- rency addresses.