GeoG2U-Bench: When Does Generation Help Understanding in Ultra-High-Resolution Remote Sensing?
Fengxiang Wang ⋅ Yueying Li ⋅ Mingshuo Chen ⋅ Boya Miao ⋅ Qiuyang Yu ⋅ Yajie Yang ⋅ Mingzhen Xu ⋅ Luqing Luo ⋅ Haonan Guo ⋅ Hongda Sun ⋅ Yulin Wang ⋅ Jun Song ⋅ Jing Zhang ⋅ Long Lan ⋅ Wenjing Yang
Abstract
Ultra-high-resolution (UHR) remote-sensing imagery captures fine-grained spatial details essential for city-scale Earth observation, yet it exposes a fundamental limitation of current multimodal large language models (MLLMs): native-resolution inputs exceed model capacity, while aggressive downsampling erases the very small objects that many questions hinge on. Generation-understanding unified multimodal models offer a promising remedy, since they can produce intermediate visual observations before answering. However, whether such capabilities genuinely improve UHR remote-sensing understanding remains an open question. Existing UHR benchmarks cannot adjudicate it: they are mostly built from mature datasets covering only a few regions, their image scales remain below real city-level satellite scenes, and they do not systematically identify which tasks benefit from intermediate visual generation. We introduce GeoG2U-Bench, a global benchmark designed to evaluate Generation-to-Understanding (G2U) in UHR remote sensing. GeoG2U-Bench spans 300 cities worldwide, with every raw satellite scene reaching $20{,}000 \times 20{,}000$ pixels, matching the scale at which UHR analysis actually occurs. It comprises 20 subtasks across five capability dimensions and roughly 3,000 expert-verified samples, each requiring models to produce auditable intermediate artifacts such as zoom-in crops, annotated local views, cross-region relation maps, and temporal rewind images. Beyond data, GeoG2U-Bench contributes a dual-protocol evaluation: Direct (answer from the raw input) versus Generate-then-Answer (generate an intermediate visual artifact, then answer), which isolates and quantifies the contribution of generation to downstream understanding. Experiments across several model families reveal that generation is not uniformly beneficial: unfaithful or off-task artifacts can actively degrade performance, whereas reliable, task-relevant artifacts yield consistent gains in cross-scale localization, multi-region comparison, temporal reasoning, and spatial layout understanding. These results show that visual generation can support UHR remote-sensing understanding, but only when it provides reliable visual evidence for the task.
Chat is not available.
Successful Page Load