Visual Harness: Grounding Multimodal Reasoning in Physics Engines
Chentao Cao ⋅ Zhanke Zhou ⋅ Sangni Duan ⋅ Bo Han ⋅ Hang Li
Abstract
Multimodal agents struggle with physical-world reasoning, particularly spatial reasoning from partial views, as the physical world is inherently complex. We think multimodal agentic reasoning about the physical world should ground its actions in deterministic feedback, rather than rely solely on internal imagination. Concretely, we propose to equip vision–language models (VLMs) with a physics engine (e.g., MuJoCo, UE5), which simulates the consequences of the agent's actions. However, granting the agent raw access to a physics engine is not sufficient, as the agent must first reconstruct the scene inside the physics engine. We therefore introduce the Visual Harness System, an orchestration system that decouples perception from reasoning with a physics engine as the backend. The perception module composes off-the-shelf perception skills, such as camera-pose estimation, open-vocabulary detection, and metric depth, into a structured scene map. The reasoning module then issues multi-turn tool calls that the engine executes over the map, and composes the final answer from the grounded feedback it receives. Experiments on spatial reasoning benchmarks show that our approach consistently outperforms strong baselines across both open- and closed-weight VLMs, with up to $16.57\%$ improvement over prior methods on the MindCube benchmark.
Chat is not available.
Successful Page Load