Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
Abstract
Generalist robot policies must ground language in visual observations, anticipate how actions change the scene, and reason about task goals. Existing VLA approaches typically address these capabilities separately: VLM-based policies emphasize semantic grounding, video or world models emphasize visual dynamics, and goal-conditioned methods emphasize future-state reasoning. This separation can make policies brittle under instruction shift and environment or dynamics perturbations. We introduce Dynin-Robotics, an omnimodal masked-diffusion VLA model that represents robot trajectories as partially observed multimodal token sequences. A single denoising backbone is trained to unify action generation, future-state and world modeling, trajectory-to-language goal understanding, and instruction-conditioned goal-state prediction. At inference time, the same model performs policy generation through block-wise action denoising, with optional goal-state prediction, and world-action joint decoding. We first present diagnostic analyses showing that VLM-based and video/world-model-based policies exhibit complementary failure modes, motivating a unified formulation. Our experiments show that Dynin-Robotics achieves competitive standard-task performance on LIBERO, while improving robustness on instruction-shift and perturbation settings in LIBERO-Plus and VLABench.