AmbiguousWorld: Benchmarking and Resolving Ambiguous Instructions in Video World Models
Abstract
Translating high-level, under-specified human commands into coherent video is fundamentally challenging. Current video generation models lack explicit reasoning capabilities and typically fail to understand ambiguous instructions, frequently resulting in physically inconsistent and causally disjointed generation. To address this, we introduce AmbiguousWorld, a novel inference-time reasoning framework that transforms video generation into a multi-modal Tree of Thoughts. We employ a Vision-Language Model as both a semantic planner to dynamically decompose fuzzy instructions into actionable subgoals, and as a closed-loop critic to prune invalid trajectories by evaluating speculative paths across multiple physical and functional dimensions. To systematically assess this, we annotate the first embodied video generation benchmark targeting ambiguous instructions based on the DROID dataset, and design a multi-dimensional VLM evaluation framework. Empirical results on state-of-the-art baselines demonstrate that our framework yields substantial improvements, boosting absolute Task Success by up to 29.0\% and the overall generative consistency weighted score by 24.2 points. Furthermore, we validate the effectiveness of our VLM evaluation framework through rigorous comparison with human assessment.