A Pre-Flight Audit Protocol for Post-Training MIA Privacy Claims on Frozen Checkpoints
Abstract
Can a large inference-time membership-inference (MIA) gap on a frozen checkpoint support a training-data exposure claim? Auditing the scoring pipeline behind a representative post-training configuration, we find three latent failures whose combined effect can inflate the target-versus-reference AUROC gap by tens of percentage points before any attack comparison is interpreted. The testbed is MIMIR arxiv (n=500 members and n=500 post-cutoff nonmembers), Min-K%++ with k=20%, Pythia-70M vs. GPT-2-medium: (P) JSONL texts scored in escaped-string form, (S) an inherited Min-K%++ sign opposite to the empirical member direction, and (F) YAML-wrapped members vs. title-prefixed nonmembers. F alone yields +42.2 percentage points (pp) from the same abstract content. After correction, 1:1 lexical matching gives +1.1 pp [-9.4,+11.3], p=0.84: no reliable target-specific residual at primary balance. WikiMIA reverses sign (-21.3 pp band); PubMed Central and GitHub fail cohort-audit checks. A wider BoW posterior band retains power and gives +14.0 pp [+7.5,+20.4] as a directional diagnostic, not a privacy conclusion. We release a 14-check verification suite and lexical-balance estimators for frozen-checkpoint MIA audits.