GeoReason: Bridging Logical Reasoning and Spatial Fidelity in Remote Sensing Segmentation
Abstract
Multimodal Large Language Models (MLLMs) have significantly advanced instruction-driven perception through the ``Embedding-as-Mask'' paradigm. However, extending this capability to Remote Sensing (RS) reasoning segmentation remains exceptionally challenging under cluttered geographical contexts and extreme scale variations, often leading to \textit{Spatial Fidelity Degradation} and \textit{Target Misidentification}. We identify a critical yet underexplored bottleneck: the sequential alienation of the segmentation token \texttt{[SEG]} from its initial visual grounding within the autoregressive generation paradigm. Produced only at the end of long reasoning chains, \texttt{[SEG]} relies on visual evidence that has been progressively diluted by linguistic abstraction and autoregressive noise. This information decay is particularly catastrophic for RS targets with minuscule pixel footprints, where even marginal fidelity loss leads to localization failure. To address these bottlenecks, we propose \textbf{GeoReason}, a novel framework designed to fortify logical deduction and preserve spatial integrity. First, we introduce an \emph{Anticipatory Prior} mechanism that injects the \texttt{[SEG]} token directly into the initial user query, shifting it from a terminal output to a primary condition. This paradigm shift transforms the token into a persistent spatial anchor, preventing information decay during long-form linguistic inference. Second, we enhance visual reasoning via a \emph{Location-aware Chain-of-Thought (CoT)} that enforces coarse-to-fine localization using bounding boxes, coupled with a \emph{Latent Patch-to-Token Alignment (PTA) Loss} that explicitly aligns the latent segmentation token with physical RS textures. Extensive experiments demonstrate that GeoReason consistently outperforms state-of-the-art methods on the Earthreason and RRSIS-D benchmarks.