BodyBench: Evaluating Adversarial Image Defenses Against AI Nudification Inpainting
Abstract
Adversarial perturbation defenses protect images from generative editing by adding imperceptible noise. Yet none of these methods have been evaluated against AI nudification, despite AI-generated non-consensual intimate imagery (AIG-NCII) being among the most documented harms of generative AI and produced largely through inpainting. We argue this is a problem-formulation failure: prior threat models do not correspond to nudification attacks, and the misalignment is structural, spanning mask shape, prompt distribution, and success criterion which causes existing evaluations to systematically overestimate protection effectiveness. We reformulate adversarial inpainting protection around this documented harm. We introduce BodyBench, a benchmark of 648 fully synthetic clothed adult subjects, each paired with 11 mask variants spanning body, face, context, and dilation regions, along with nudification and redressing attack prompts. We propose three harm-aligned metrics: sexualization, naturalness, and recognizability that jointly determine whether an inpainted output constitutes a successful nudification attack. Auditing four recent protection methods (PhotoGuard, DiffusionGuard, AdvPaint, DiffVax) on BodyBench, we find that protection succeeds in only one configuration of one method, and only under exact mask alignment between defender and attacker---a condition no real defender can guarantee and one that we show is easily always bypassed. We also show that standard metrics such as PSNR are essentially uncorrelated with nudification protection. BodyBench makes this misalignment measurable and provides a foundation for defenses grounded in documented harm.