DualWorldBench: Can Agents Plan Deliveries across Symbolic and Grounded Worlds?
Abstract
Recent multimodal agents have shown promising progress in embodied planning and web interaction. However, existing evaluations often rely on simplified tasks with limited combinatorial structure, leaving it unclear whether current agents can robustly integrate perception and reasoning for multi-goal planning under temporal, physical, and multimodal constraints. We introduce DualWorldBench, a unified benchmark for evaluating agent planning under controlled symbolic--temporal and grounded--topological complexity. Built around a shared pickup--delivery interface, DualWorldBench consists of three complementary components: TextWorld, which isolates symbolic planning under varying task compositions and release dynamics; SimWorld, which evaluates grounded path planning under controlled topological variation with text, image, or multimodal inputs; and DualWorld-VQA, a diagnostic suite spanning perception, scene understanding, and planning. Experiments show that current state-of-the-art agents remain brittle under increasing combinatorial complexity. Multimodal input does not consistently improve planning and can even degrade performance, revealing difficulties in aligning symbolic task specifications with grounded topology. Diagnostics further show that models handle perception or understanding in isolation, but fail when decision-making requires integrating both. These findings establish DualWorldBench as a challenging testbed for agents with tighter perception--reasoning--planning integration. All datasets and code are publicly available at our project website https://dualworld.netlify.app/.