See to Believe: Segmentation by Reasoning with Visual Evidence
Abstract
Reasoning segmentation requires localizing objects from complex, implicit textual queries. Existing methods either rely on the MLLM to reason implicitly within hidden representations, or produce explicit but purely textual chain-of-thought that terminates in a coarse bounding box. In both cases, the reasoning process is not grounded in visual evidence at each step, leaving intermediate errors unverifiable, which can propagate to the final prediction. We propose Reasoning with Visual Anchors (ReVA), a framework where each chain-of-thought step is explicitly grounded to a specific image region. ReVA introduces anchor tokens into intermediate reasoning steps, each directly aligned with region-level features in the image encoder's latent space. As the model reasons step by step, each anchor attends to its corresponding image region, grounding inference in concrete visual evidence. To supervise the full reasoning chain, we further require each anchor to correctly transition to the next along the spatial relation described in the text. This progressively guides the model toward the target and produces the correct segmentation mask. To support training, we construct a large-scale Chain-of-Anchor dataset of over 245K quality-filtered samples from the RefCOCO series, each annotated with a spatially grounded chain-of-thought containing explicit anchor regions. The dataset is constructed by leveraging a strong MLLM to generate reasoning traces, followed by a three-stage verification pipeline ensuring spatial accuracy and reasoning quality. Experiments on reasoning and referring segmentation benchmarks demonstrate that ReVA achieves state-of-the-art performance (e.g., +5.4 gIoU and +6.7 cIoU on the ReasonSeg test set), with ablations showing that reasoning in visual evidence is key to the improvement.