G-PAC: Constructing Cohesive Pseudo-Features for Generalizable Physical Adversarial Camouflage
Abstract
Physical-domain adversarial attacks present critical security threats to diverse AI systems, notably autonomous driving. Generalization is an inherent requirement and a primary obstacle in physical attacks, manifesting as the ability of adversarial patterns to persist across varying environmental configurations and to transfer to unseen victim models. While prior methods demonstrate generalization across specific models, their performance drops significantly on open-vocabulary foundation models, with some approaches becoming entirely ineffective. This lack of generalization stems from the failure of adversarial camouflage to construct a cohesive pseudo-feature across perspectives, lacking the stable semantic representation characteristic of natural objects. Such a deficiency arises because existing attack paradigms rely on independent view-level optimization and non-targeted objectives, which lead to semantically scattered features and inconsistent optimization directions. To address these limitations, we propose Generalizable Physical Adversarial Camouflage (G-PAC), a framework that transitions from view-specific suppression to coordinated object-level representation learning. G-PAC maintains a momentum-updated global pseudo-feature center as a stable adversarial anchor to guide optimization across diverse configurations. By leveraging the feature space of a self-supervised foundation model as a semantic prior, we introduce a contrastive regularization term to encourage feature aggregation while ensuring adversarial potency. Comprehensive digital and physical evaluations demonstrate the effectiveness of G-PAC. Notably, it outperforms the strongest baseline by an average AP@0.5 margin of 0.10 across seven diverse detector architectures in digital settings, and further extends this margin to 0.12 in real-world physical tests. Furthermore, G-PAC exhibits profound black-box transferability against highly resilient open-vocabulary foundation models, inducing an additional absolute AP@0.5 reduction of 0.30 on GLIP.