GeoX: Mastering Geospatial Reasoning Through Self-Play and Verifiable Rewards
Abstract
Geospatial reasoning requires solving image-grounded problems over complex object relationships and scene structures. However, developing this capability is hindered by the cost of annotating a vast and combinatorial question space. We propose GeoX, a self-play framework that acquires spatial logic through executable programs and verifiable rewards in the absence of large-scale human-curated data. Given a satellite or aerial image, our framework employs a single multimodal policy to alternate between proposing spatial problems with executable programs and solving them under three reasoning modes of abduction, deduction, and induction, using a segmentation tool and standard computational libraries. Program execution turns each program into a reward that jointly optimizes the two roles via reinforcement learning. GeoX consistently improves its base VLMs by up to 5.5 points on average, matching or exceeding conventional baselines trained on millions of curated samples. Alongside the framework, we release a benchmark for geospatial reasoning curated from self-play generations.