EviSAM3: Evidence-Driven SAM3 for Referring Remote Sensing Image Segmentation
Abstract
Referring Remote Sensing Image Segmentation (RRSIS) aims to segment target objects from remote sensing images based on natural language descriptions, requiring precise vision–language alignment under significant ambiguity. Despite the strong generalization ability of Segment Anything Model 3 (SAM3), the large spatial extent and varying scales of targets in remote sensing images intensify the ambiguity of referring expressions. This poses a major challenge for fine-grained segmentation. To address the challenge, we propose EviSAM3, an evidence-driven parameter-efficient adaptation framework that mimics human decision-making by dynamically accumulating partial evidence. Specifically, EviSAM3 progressively integrates cross-modal evidence through an evidence memory bank, where each incoming piece is aggregated via momentum-based updates. As evidence accumulates, the model performs iterative evidence rectification to resolve ambiguity under incomplete observations. Meanwhile, the contributions of different evidence are adaptively modulated based on their diagnostic relevance, allowing more informative evidence to dominate the reasoning process and guide the final prediction. Extensive experiments demonstrate that EviSAM3 consistently outperforms state-of-the-art methods, particularly in challenging scenarios with high ambiguity and significant scale variation. The code is available in the supplementary material.