Learning Data-free Universal Adversarial Perturbation with Hybrid Priors and Gradient-Guided Sharpness Regularization
Abstract
Universal adversarial perturbations (UAPs) aim to learn a single input-agnostic perturbation that consistently induces misclassification across diverse inputs, exposing systemic vulnerabilities of deep models. To enhance practical applicability, data-free UAP methods replace real data with samples drawn from handcrafted synthetic priors and optimize UAPs via activation maximization. However, existing approaches rely on a single synthetic prior throughout training, which biases optimization toward a narrow feature subspace and necessitates costly per-prior tuning. Moreover, the commonly used stochastic gradient ascent (SGA) updates are inherently unstable and tend to converge to sharp regions of the loss landscape, weakening black-box transferability. To overcome these limitations, we propose HPSR-UAP, a data-free framework that couples hybrid-prior reweighting with gradient-guided sharpness regularization. Specifically, we develop a dynamic reweighting mechanism over multiple synthetic priors to mitigate prior bias and avoid overfitting to any single distribution. Building upon this, we introduce a gradient-guided sharpness regularization that smooths the prior-weighted activation loss landscape within a local perturbation neighborhood, promoting updates towards flatter and more transferable directions. Extensive experiments on ImageNet demonstrate that HPSR-UAP outperforms state-of-the-art methods in both white-box and black-box data-free attack settings. We further introduce a multi-granularity block-wise similarity metric that quantifies spatial repetitiveness within UAPs, revealing a strong correlation between structural regularity and attack performance.