eSAM: Editing SAM3 Attention for Training-Free Referring Segmentation
Yaoting Wang ⋅ Yun Zhou ⋅ Hengrui Hu ⋅ Chang Liu ⋅ Henghui Ding
Abstract
Referring expression segmentation (RES) aims to produce a pixel-level mask for the object described by a free-form natural language expression, in either images (RIS) or videos (RVOS). Existing zero-shot approaches reduce the cost of mask-language annotation but still rely on auxiliary training stages or multi-model pipelines. The recently released SAM 3 introduces native text-prompted segmentation within a single foundation model. However, directly applying it to RES generates unexpected weak performance. We trace this gap to an attention sink in SAM 3's multimodal fusion encoder, where the start-of-text token absorbs on average $54.8\%$ of each image patch's cross-attention mass, leaving content tokens with limited influence on the fused visual features. Suppressing the sink alone is insufficient, as the released attention mass spreads without direction; closing the gap requires both suppressing the sink and providing patch-dependent spatial guidance that routes attention to semantically relevant content tokens. Building on these findings, we introduce eSAM, a fully training-free framework that edits SAM~3's cross-attention without any parameter updates. With Attention Mask Editing (AME), eSAM edits the attention map, suppressing the sink while injecting a CLIPSeg-derived spatial prior that directs released attention toward relevant content tokens. Further with Value Tensor Editing (VTE), eSAM rescales the value tensors, amplifying the magnitude with which content tokens contribute to the fused output. On both RIS (RefCOCO/+/g) and RVOS (Ref-DAVIS17, Ref-YouTube-VOS, MeViS) benchmarks, eSAM achieves new state-of-the-art results among training-free methods.
Chat is not available.
Successful Page Load