Evaluating an Evaluation: Membership Inference Attacks as Machine Unlearning Diagnostics
Umid Suleymanov ⋅ Laman Aliyeva ⋅ Nihat Abdullayev ⋅ Saida Zarbiyeva ⋅ Murat Kantarcioglu
Abstract
Evaluation practices in machine learning are increasingly objects of scientific study in their own right: their assumptions, statistical properties, and failure modes determine which research conclusions can be trusted. We audit one such practice - the use of aggregate-subset Membership Inference Attacks (MIAs) for evaluating machine unlearning - across 8 dataset-architecture pairs, 4 unlearning methods, and over 500 experimental runs, identifying five failure modes: (1) limited and regime-dependent identifiability between baseline and retrained models, with KS-test indistinguishability ($p > 0.05$) in $\approx$42\% of experiment--task pairs; (2) regime-dependent signal-to-noise collapse; (3) monotonicity violations in 5 of 8 experiments under graded forgetting; (4) inter-task rank inconsistency, and (5) regime confounding by model training quality. To quantify these failures, we develop the Membership Inference Attack Unlearning Score (MIAU), a normalized diagnostic framework integrating three MIA comparisons as gap closure fractions between baseline and retrained references. We formalize four validity properties any unlearning metric should satisfy. Our findings indicate that aggregate-subset MIAs are reliable unlearning diagnostics only in narrow high-memorization regimes, and motivate complementary evaluation paradigms, alongside MIA-based protocols.
Chat is not available.
Successful Page Load