When the Attacker Is the Metric: Auditing Re-identification Evaluation for Learned Data Sanitizers
Abstract
Learned sanitizers suppress biometric cues in chest radiographs, but their measured privacy depends on the attacker used to evaluate them. On unsanitized radiographs, the evaluator reproduced from the released code links only 4.07% of same-patient pairs at 1% FPR (79.92 AUC), while the same architecture initialized from the reference patient verifier reaches 87.05% (99.45 AUC)—showing that low measured leakage can reflect a weak evaluator rather than an effective sanitizer. Auditing this evaluation revises reported privacy numbers substantially, but does not overturn the original qualitative conclusion: under the strongest evaluated attacker, the learned field still leaks less than utility-matched blur and random elastic controls. Using PriCheXy-Net as a case study, our audit protocol checks attacker adequacy across 11 configurations. Under the released training protocol, the frozen verifier measures more linkage than the ImageNet-initialized evaluator in every configuration. In a separate exploratory three-seed extension, matching training and validation pairing to deployment raises learned-field/verifier-init AUC from 81.61 to 94.74 and TPR at 1% FPR from 4.81% to 27.29%. Patient-clustered uncertainty quantifies utility differences without establishing equivalence. These results, limited to one learned sanitizer and one patient-level dataset, show how attacker provenance and training/validation pairing shape empirical privacy estimates.