Can Model Similarity Verify a Machine Unlearning Action? Evidence from Cross-Auditor and Retraining Controls
Abstract
We show that a favourable model measurement cannot identify which deletion action produced a model. A prospectively frozen Fashion-MNIST study crossed 20 training seeds, five candidate constructions, and three independently implemented audit families. On the 60 non-anchor case–seed units, the applicable parameter and predictive auditors disagreed 39 times (65%). A loss-only membership auditor abstained in all 100 units because it lacked the frozen anchor discrimination required to decide. Predictive Jensen–Shannon divergence passed wrong-scope retraining once in 20 trials at the frozen 0.10 cutoff (5%; exact 95% interval, 0.13–24.87%); the pass disappears at 0.05. A separately frozen control showed distributional overlap: valid same-scope and wrong-scope retraining each produced 1 PASS, 10 UNKNOWN, and 9 FAIL predictive decisions, and wrong-scope retraining was closer to the single oracle in 9 of 20 paired seeds. These measurements answer questions about model states. Evidence for an action also requires a binding among the request, resolved data scope, execution, and resulting state.