Attack-Inference Aligned Universal Adversarial Attacks on Video Object Segmentation
Abstract
Modern video object segmentation (VOS) relies on a two-stage paradigm for robust performance: target initialization and memory-based propagation. However, existing adversarial attacks on VOS often overlook this fundamental feature. To bridge this gap, we propose AUV, an adversarial framework that Aligns with the Universal two-stage paradigm of VOS models. Specifically, AUV optimizes a single universal adversarial perturbation (UAP) through decoupled training strategies tailored for both initialization and propagation stages. To ensure prompt-agnosticism across all modalities, including masks, points, boxes, and language, AUV optimizes the UAP to disrupt the shared memory space for latent-level target erosion. Extensive experiments show that AUV establishes a new state-of-the-art for VOS attacks, such as degrading mask-based SAM3 to only 8.4 J&F on DAVIS17-val (compared to 55.4 for previous SOTA under the same training setup), while consistently outperforming prior methods on other VOS models including SAM2, Cutie, and XMem. Notably, AUV exhibits superior temporal dynamics with the fastest onset and the most enduring persistence, while uniquely supporting flexible any-time mid-video attack, revealing vulnerabilities in real-world streaming applications. Our code and checkpoint will be made publicly available.