StablePrivacy: Zero-Shot Full-Body Video Anonymization Using Diffusion Models For Downstream Data Utility
Abstract
Releasing visual datasets containing people requires removing identity information without destroying the signal that makes the data useful for training. We present StablePrivacy, a framework for full-body video anonymization that runs entirely at inference time: it uses a pretrained diffusion model with no fine-tuning or dataset-specific training, and sanitizes a training set before release. It preserves pose and scene context via ControlNet conditioning and latent-space inpainting, a memory bank enforces temporal consistency, and an injection module retains task-relevant structure. Measuring utility by training downstream models on anonymized data and testing on original data across video instance segmentation, action recognition and re-identification, StablePrivacy attains a stronger privacy--utility trade-off than existing methods. We further find that a stronger, higher-fidelity generator does not by itself yield better utility, indicating the trade-off is governed by the anonymization design rather than the generator's capacity.