When to Inject the Target: Stage-Decoupled Guidance for Diffusion-Based Targeted Adversarial Attacks
Abstract
Diffusion-based generative attacks have emerged as a promising paradigm for targeted adversarial example generation, leveraging generative priors to improve transferability and visual fidelity. However, existing methods typically entangle image reconstruction and target manipulation within the same reverse denoising trajectory, where target guidance can disrupt early structural recovery. This coupling creates a persistent trade-off among attack effectiveness, transformation robustness, and perceptual fidelity. In this paper, we propose SDTI (Stage Decoupled Target Injection), a stage-decoupled reverse diffusion framework that separates source-structure preservation from adversarial target injection along the denoising timeline. SDTI first maps the input image to an intermediate latent through null-text DDIM inversion, then recovers and anchors the source structure with a frozen null-text denoising prefix, and finally activates target-conditioned LoRA guidance only in the late denoising suffix to inject adversarial target semantics. To confine adversarial optimization to this suffix phase, we further introduce suffix-wise supervision with inter-step gradient truncation. Extensive experiments on ImageNet-NeurIPS, robust victim models, common input transformations, and MS-COCO cross-domain transfer show that SDTI achieves a favorable attack--fidelity--robustness trade-off, improving transformation robustness and perceptual quality while maintaining competitive targeted transferability. The code is available at: https://anonymous.4open.science/r/AY_code-1431.