Localizing and Repairing Sparse-Prompt Failure in SAM Decoders via Box-to-Point Counterfactual
Abstract
Promptable segmentation foundation models such as SAM and its variants achieve strong performance with informative prompts, yet exhibit a striking prompt asymmetry where dense prompts like bounding boxes yield accurate segmentation, whereas sparse point prompts substantially degrade performance. This performance gap is widely observed, yet its internal decoder mechanism remains underexplored. In this work, we revisit the failure of point prompts from a causal perspective. By adapting activation patching to SAM-style decoders with a token-aligned intervention scheme, we localize the earliest recoverable divergence to a single early operation, i.e., the first decoder image-to-token cross-attention update, where prompt-conditioned information first enters the image stream. A sparse autoencoder further reveals that this mediation is concentrated in approximately 50 feature directions out of 4,096, and reverse corruption confirms that the identified layer is the entry point of a propagated image-stream route rather than a locally sufficient state. Guided by this localization, we propose the Causal Bottleneck Adapter (CBA), a lightweight residual module placed at the identified entry layer with the base model frozen. Evaluated on 97,996 medical image-target pairs across CT, MRI, ultrasound, and endoscopy, and validated across various SAM-family models, CBA closes 0.891 of the box-to-point gap with fewer than 9K trainable parameters, outperforming prior repair baselines up to 50× larger.