BODYBENCH: Evaluating Adversarial Image Defenses Against AI Nudification Inpainting
Abstract
AI-generated non-consensual intimate imagery (AIG-NCII) is the sexualized or nude depiction of a person without their consent. It is currently produced at scale through commercial applications and open-source workflows, many of which are created using diffusion inpainting methods. Adversarial perturbations is one of the very few proactive technical defenses available to disrupt unwanted inpainting-based editing. In theory, a subject may add imperceptible noise to a photo before publishing it, and later generative edits on this photo will fail. A growing literature has developed these methods. However, these methods neither target AI nudification nor evaluate against its specific threat model. This is a failure in threat-modeling, as a method’s choice of inpainting mask and prompt together defines an implied threat model. Prior work assumes benign edits–such as removing or adding an object, or background change, and evaluates success with generic metrics such as PSNR and SSIM. Nudification takes on the opposite structure. The attacker regenerates the body area, removing what was once covered by clothing and replacing it with a nude body, while preserving the face, identity, and background context. This means that prior methods are misaligned to the AI-nudification task. While existing defenses may appear to be effective against generic inpainting tasks, we show they fail to prevent real-life AIG-NCII harm.