DOME: Drift-Adaptive On-Policy Motion Erasure in Video Diffusion Transformers
Abstract
Concept erasure works well for static visual attributes in text-to-image diffusion, but does not transfer to motion concept erasure in video diffusion transformers. We trace this failure to an off-policy trajectory mismatch: once weights are edited, the generation trajectory followed by the edited model deviates from the original trajectory used for supervision, so state-local velocity matching fails to control the motion patterns realized at inference. As a consequence of this mismatch, off-policy weight modification can drive training loss to convergence while the target motion remains fully present in generated videos. We formalize this gap through a Gronwall-type bound showing how weight-induced velocity perturbations propagate into trajectory divergence, and a residual-error bound showing that off-policy training, even at perfect convergence, leaves a nonzero erasure error on the edited trajectory. To address this problem, we introduce {DOME (Drift-Adaptive On-Policy Motion Erasure)}, a trajectory-aligned framework for motion concept erasure in video diffusion transformers. DOME trains on states from the edited model's own trajectory and applies \emph{drift-adaptive anticipation}, which shifts supervision along the observed divergence between the edited and original trajectories, turning trajectory drift from a failure symptom into a training signal. DOME edits the text-conditioned pathway via cross-attention LoRA, and the learned adapter can be merged into the base model so that inference incurs no additional overhead. On Wan~2.1-T2V across 20 motion concepts, DOME consistently suppresses localized contact actions while showing limited effect on whole-body dynamics, validated through both motion-consistency scoring and an independent action classifier. Systematic ablations show that drift-adaptive anticipation, not on-policy rollout alone, is the critical component distinguishing DOME from off-policy and standard on-policy training.