Seed-Your-Motion: Householder Orthogonal Noise for Motion-Controllable Video Diffusion Models
Abstract
Diffusion-based video generation has achieved impressive visual fidelity, yet enforcing temporally consistent motion under explicit control signals remains challenging. Recent noise-warping approaches construct temporally correlated noise by transporting the initial latent noise via diffeomorphic deformation fields derived from optical flow, but this spatial resampling inherently destroys the i.i.d. Gaussian structure of the noise, requiring post-hoc correction steps at the cost of non-trivial computational overhead. We propose \textbf{Seed-Your-Motion}, by conditioning the a priori latent noise on optical flow via per-location orthogonal transformations in latent channel space, parameterized as products of the Householder Transformation. This channel-space formulation has an advantage of preserving the spatial Gaussianity exactly by construction, requiring neither Jacobian corrections nor post-hoc distribution restoration, while at the same time inducing structured cross-frame correlations that ensures geometric consistency across time. With no learnable parameters and a single forward pass, our method is agnostic to model architecture and can be plugged seamlessly into any video diffusion fine-tuning pipeline. Extensive experiments on motion-transfer and camera-controlled video generation benchmarks demonstrate that our method matches or outperform the state-of-the-art noise-warping baseline, while achieving a 2.7× speedup in processing time. The results suggest that the channel-space Householder transformations provide a principled and efficient alternative to noise-warping based motion-controllable video diffusions.