Motion Forcing: Decoupling Ego and Object Motion via Sparse Inputs for Structured Video Generation
Abstract
Structured video generation (e.g., autonomous driving) faces a fundamental domain gap between simple human control and dense physical reality. Text-to-Video models accept accessible inputs but struggle with precise spatial guidance, while dedicated autonomous driving world models require expensive, dense 3D annotations that are difficult for users to provide. To bridge this gap, we introduce \textbf{Motion Forcing}, a structured video generation framework that achieves precise multi-agent control using extremely sparse 2D inputs. Our key insight is to explicitly bridge human intent and visual synthesis via a hierarchical \textbf{``Point-Shape-Appearance''} paradigm. This approach decomposes generation into verifiable stages: modeling user intent as sparse geometric points, expanding them into dynamic depth maps to explicitly resolve 3D geometry and decouple ego-motion from object dynamics, and finally rendering high-fidelity textures. Furthermore, to elevate the model from passive instruction-following to active physical reasoning, we employ a \textbf{Masked Point Recovery} strategy. By forcing the reconstruction of dynamic depth from occluded trajectories, the model internalizes latent physical laws, enabling causal inference. Extensive experiments demonstrate that Motion Forcing significantly outperforms state-of-the-art baselines on large-scale autonomous driving benchmarks. Additional validation in rigid-body physics and robotic manipulation confirms the structural integrity and generalizability of our framework. \textbf{To facilitate future research, we have released our model weights and code.}