Controllable Dynamic 3D Shape Generation via 3D Trajectories and Text
Abstract
We present T2Mo, the first feed-forward framework for controllable dynamic 3D shape generation of general objects conditioned on 3D trajectories and text prompts. Existing text-driven methods struggle to specify precise motion due to the ambiguity of natural language. To address this, we introduce 3D trajectories as a direct spatial control signal and propose a shape-grounded trajectory embedding that maps arbitrary user-provided trajectories into geometry-aware condition tokens aligned with the input shape. Injecting these tokens into the generative backbone enables highly controllable motion generation while handling varying trajectory numbers and distributions. We conduct extensive comparisons against text-based baselines and trajectory-guided video generation workarounds. Quantitative and qualitative evaluations, along with user studies, show that our method produces motions that more faithfully follow the given prompts with higher expressiveness while preserving motion quality. Our code and weights will be publicly released.