Robust Concept Unlearning in Diffusion Models via Directional Stability Regularization
Abstract
Post-hoc concept unlearning is a practical approach to remove undesired content from text-to-image diffusion models. However, existing methods leave unlearned models fragile: small perturbations to prompts, text embeddings, or sampling-time latent states can recover supposedly erased content, even when standard endpoint metrics suggest successful erasure under nominal prompts. We study this fragility through the stability of the denoising trajectory of diffusion models. Our key observation is that prompt-, embedding-, and sampling-based attacks share a similar mechanism of perturbing the sampling dynamics, and erased concepts can reappear when the unlearned model amplifies these concept-relevant perturbations. We formalize this view with a finite-time stability analysis and derive a tractable directional sharpness measure along a concept-relevant direction. Motivated by this analysis, we propose Directional Stability Regularization (DSR), a plug-in regularizer that discourages expansion in erased-concept directions without directly penalizing unrelated directions. DSR is compatible with classical score-based unlearning objectives, requires no adversarial training or attack-specific inner loop, adds no inference-time overhead, and can be estimated efficiently with Jacobian-vector products. Across the evaluated erasure settings, DSR reduces recovery under prompt-, embedding-, and sampling-based attacks while largely preserving prompt alignment and image quality metrics.