GeoMIND: A Benchmark for Spatial Understanding in Robotic Manipulation
Abstract
Recent advances in robotic manipulation have made rapid progress in mapping visual observations and language instructions to actions. However, reliable manipulation in open-world environments requires robust spatial understanding, the ability to identify targets from complex layouts (spatial reasoning) and adapt to scenes that continuously evolve through interaction (spatial memory). Yet existing embodied benchmarks fall short in evaluating spatial understanding: they often introduce strong visual cues that bypass spatial reasoning, and focus on static scenes that offer little insight into an agent’s capacity for spatial memory. To address this gap, we introduce GeoMind, a systematic benchmark for structured spatial understanding through tabletop manipulation. GeoMind comprises 54 evaluated task variants derived from 27 canonical task designs organized around spatial reasoning and spatial memory. It also incorporates a scalable pipeline that automates the generation and annotation process of task instances with verified targets and executable interactions. Extensive evaluations across diverse embodied agents reveal their consistent weaknesses in spatial understanding. Our findings highlight that these agents struggle to track identities after spatial changes and convert relational or geometric constraints into executable actions.