DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control
Teli Ma ⋅ Jia Zheng ⋅ Zifan Wang ⋅ Chunli Jiang ⋅ Andy Cui ⋅ Junwei Liang ⋅ Shuo Yang
Abstract
Vision-Language-Action (VLA) models inherit their visual backbone from static image-text pretraining, leaving physical dynamics to be learned from scarce action data. Generative video models, by contrast, already encode motion, contact, and implicit physics at internet scale. We introduce DiT4DiT, a Video-Action Model (VAM) that couples a Video Diffusion Transformer with an Action Diffusion Transformer and co-trains them under a dual flow-matching objective with a tri-timestep scheme for video supervision, action supervision, and feature extraction. The key insight is that the policy does not need to generate a future video: we intercept the Video DiT's hidden state at a single fixed flow timestep, turning a multi-step video rollout into a one-shot feature extraction consumed by the Action DiT. Across benchmarks, DiT4DiT sets new SOTA averages on LIBERO (98.6%) and the 24-task RoboCasa-GR1 suite (56.7%), outperforming strong VLA and video-based baselines. On a real Unitree G1, DiT4DiT reaches 73.8% on eight tabletop tasks and 72.2% on three whole-body loco-manipulation tasks, running up to 12$\times$ faster than prior VAM baselines while also exceeding them in accuracy. To our knowledge, this is the first video-generation-based policy to run humanoid whole-body control at real-time rates. It further generalizes zero-shot to unseen object categories and scene variations. Together, these results indicate that video-dynamics priors, accessed through a single forward pass, are a practical and scalable foundation for generalist robot policies.
Chat is not available.
Successful Page Load