MoSync: 3D Motion Synthesis from End-Point Conditioning
Abstract
Generating temporally coherent 3D motion from static scenes with precise control remains challenging. Static 3D editing lacks temporal dynamics, whereas 4D generation offers limited localized control and often suffers from geometric distortion, view inconsistency, and background drift. To address these limitations, we present MoSync, a progressive multi-stage framework for controllable 3D motion synthesis from 3D Gaussian Splatting using user-specified 3D drag points and spatial masks. Our method first injects sparse drag controls and masks into a pre-trained video diffusion model to establish a coarse spatio-temporal deformation field, applying geometry-preserving constraints and selective local optimization to increase optimization stability and maintain structural integrity. Next, score distillation refines fine visual details, while a decoupled offset parameterization isolates and eliminates unintended background drifts. Extensive experiments demonstrate that our framework achieves superior spatio-temporal coherence, plausible motion, sharper details, and higher-fidelity novel-view rendering compared to existing baselines.