Exploring Advertising Manipulation in Diffusion Image Generation
Abstract
As text-to-image diffusion models (T2I DMs) become widely deployed, adversarial advertising has emerged as a realistic threat: an attacker may compromise a T2I DM so that it implants target product brands into images generated from users' non-advertising prompts. Two key challenges remain largely unresolved in this setting: achieving natural and semantically coherent adversarial advertisement and ensuring robust adversarial advertisement. To address these challenges, we develop a new estimation algorithm for the multivariate continuously scaled phase-type with Lévy (MCPHL) distribution to capture the intrinsic distribution of natural advertisement prompts in the prompt-embedding space. With the estimated MCPHL, we construct an attack that pushes non-advertising prompts toward high-density regions of this distribution, making the resulting perturbed prompts better aligned with natural advertising prompts. We then propose a masked parameter smoothing approach grounded in mollification theory, yielding a smoothed T2I DM equipped with a dimension-invariant certified guarantee against advertisement degradation under model fine-tuning. The masking mechanism preserves utility by avoiding unnecessary smoothing on sensitive parameters, and our theoretical analysis shows that the smoothed model better preserves adversarial advertisements under fine-tuning, while maintaining better generation quality.