Points That Matter: Informative Point Selection for Training-Free Referring Segmentation
Pouya Sadeghi ⋅ Xiangnan He ⋅ Pablo Guerrero ⋅ C Thomas ⋅ Alexander Wong ⋅ Sirisha Rambhatla
Abstract
A small set of point prompts can connect language-based localization to pixel-level segmentation, but its usefulness depends on which points are selected. A coarse bounding box may include distractors and background, and poorly chosen points can further misdirect the segmenter. We study how to construct informative prompts under a small point budget without task-specific training. Our method, PinPoint, fuses four classical visual cues to select spatially diverse interior points, uses a frozen vision-language model (VLM) to verify their positive or negative semantics, and filters low-confidence labels before prompting SAM. Under a matched five-point budget with boxes, verifier, filter, and segmenter fixed, PinPoint improves cumulative Intersection-over-Union (cIoU) by $5.2$--$10.1$ points over the strongest of five non-uniform heuristic selectors on each RefCOCO split, and by $11.9$--$18.2$ points over uniform sampling across RefCOCO/+/g. On gRefCOCO, it improves masks for returned boxes while leaving no-target accuracy unchanged. The gains come from improved point prompts, with both foundation models kept frozen.
Chat is not available.
Successful Page Load