Generation-for-Understanding with Structured Action Scripts for Embodied Multimodal Learning
Abstract
Vision-language models (VLMs) are increasingly used as perception and reasoning modules for embodied intelligence, yet they remain brittle when understanding fine-grained physical processes in videos. While existing models often recognize the overall task, they frequently fail to recover the intermediate actions and result states that determine how an embodied interaction actually unfolds. We argue that this limitation stems in part from a supervision bottleneck: conventional video-language data is typically constructed through post-hoc annotation, where captions are written after video collection and often provide only coarse, loosely aligned descriptions of dynamic physical events. We propose Gen4Understand, a generation-for-understanding framework that reverses this data flow. Instead of captioning videos after they are observed, Gen4Understand first represents embodied interactions as structured action scripts composed of atomic events and state deltas, then synthesizes videos conditioned on these scripts and real start-end frame anchors. A verifier-calibrator scores each generated candidate using its conditioning script, endpoint anchors, and video content, filtering for semantic consistency, endpoint fidelity, physical plausibility, temporal coherence, and visual quality. The retained atomic clips are further merged into action-, subtask-, and task-level supervision and converted into captioning and question-answering data for VLM adaptation. Experiments across public embodied video splits, distractor-rich hallucination tests, simulated VLA benchmarks, and real-robot tasks show that Gen4Understand consistently improves fine-grained action understanding, result-state recognition, temporal ordering, hallucination resistance, and downstream embodied control. These results suggest that controllable video generation can serve not merely as data augmentation, but as a mechanism for creating calibrated process-level supervision for embodied intelligence.