Are Well-Trained Surrogates Optimal? Rethinking the Surrogate Role with Instability for Unlearnable Examples
Abstract
Unlearnable Examples (UEs) aim to protect data from unauthorized model training by injecting imperceptible perturbations that degrade model generalization. Existing methods typically rely on well-trained surrogate models to generate such perturbations, implicitly assuming that stronger surrogates yield better degradation. In this work, we challenge this widely adopted assumption and show that it is fundamentally suboptimal. We provide theoretical insights by establishing an upper bound on the target model's test loss, showing that perturbation effectiveness is closely related to the surrogate’s sensitivity to input perturbations. Specifically, a less robust surrogate model leads to stronger poisoning performance. Motivated by this analysis, we propose a simple yet effective stability metric to select optimal surrogate models for single-level methods. We further extend this principle to bi-level optimization frameworks. Beyond revealing the bi-level method's implicit reliance on non-robust surrogates, we substantially boost its poisoning performance by the stability-based surrogate selection as well. Experiments on five datasets and four representative UE methods show that our selected surrogates reduce the target model's test accuracy by an average of 18\% compared with well-trained surrogates. Notably, this is the first work to reduce CIFAR-10 test accuracy to 10\% under a perturbation budget of 1/255, whereas prior UE methods typically rely on 8/255.