Reliable Visual Grounding Beyond Mask Overlap: Controlled Interventions and Output Repair for Referring Segmentation
Abstract
Visual grounding systems are commonly rewarded for spatial overlap even when a language instruction specifies more than object extent, such as relations, multiplicity, negation, or an unavailable referent. We study this gap in referring-expression segmentation using RefSeg-CA, a controlled reliability protocol with 2,400 diagnostic records, 1,800 same-scene intervention pairs, and a positive-only RefCOCO control of 1,000 expressions with official COCO masks. Nine frozen configurations span dense segmenters, generalized referring segmentation, detector-to-SAM pipelines, multimodal localization interfaces, and direct-box decoder ablations. Positive-target mask quality does not determine exact grounded identity: across the nine frozen systems, the descriptive Spearman association between positive-target IoU and exact target-set accuracy is ρ = −0.117. We therefore evaluate grounded response validity explicitly and test a training-free instruction-conditioned mask resolver. The resolver raises macro exact target-set accuracy from 0.369 to 0.553 and same-scene pair correctness at IoU 0.5 from 0.208 to 0.313, while mean RefCOCO IoU changes only from 0.380 to 0.374. A cross-fitted abstention layer improves controlled empty-response compliance to 0.950 but reduces natural-image IoU to 0.245. These results show that reliable visual grounding should evaluate what is selected, whether a response is warranted, and whether predictions change correctly under controlled linguistic interventions—not only where foreground pixels overlap.