ONE-SHOT: Compositional Human-Environment Video Synthesis via Spatial-Decoupled Motion Injection and Hybrid Context Integration
Abstract
Recent advances in Video Foundation Models (VFMs) have revolutionized human-centric video synthesis, yet fine-grained and independent control of subjects and scenes remains a critical challenge. Recent attempts to achieve joint human-environment control often rely on explicit human-scene 3D alignment, improving geometric control but sacrificing generative flexibility and long-horizon consistency while requiring heavy 3D pre-processing. We present ONE-SHOT, a parameter-efficient framework for compositional human-environment video generation. Our key insight is to factorize the generative process into disentangled signals, separating human dynamics from scene cues. Specifically, our canonical-space motion injection and Dynamic-Grounded-RoPE map canonical human motion to the target video region, enabling accurate motion and placement control without heuristic human-scene 3D alignment. Hybrid Context Integration further maintains subject and scene consistency for minute-level synthesis. Experiments show that ONE-SHOT consistently outperforms state-of-the-art methods in structural control, flexible composition, text-guided semantic control, and long-horizon consistency.