SurgReasoner: Surgical Reasoning Segmentation with Dynamic Difficulty-Aware Reinforcement Learning
Abstract
Surgical reasoning segmentation requires a model to infer the target instrument or anatomical structure from an implicit clinical query and produce a pixel-level mask. This setting is far more challenging than category-level surgical segmentation, as targets are specified by functional role, spatial relation, or anatomical interaction rather than explicit class names. It also exposes a key limitation of existing reinforcement learning-based visual grounding methods: standard reward normalization tends to under-optimize clinically important hard samples, such as small instruments, occluded organs, and ambiguous targets. To address this, we introduce SurgSeg-14K, a surgical reasoning segmentation benchmark with over 14K image-mask pairs across seven surgical scenarios and 30 fine-grained categories. Each sample links an implicit clinical query to a binary mask, spatial annotations, and a reasoning trajectory. We further propose SurgReasoner, which decomposes the task into two stages: clinical query interpretation for spatial prompt generation, and prompt-based mask generation with a frozen segmentation model. To improve reinforcement learning for spatial reasoning, we introduce a dynamic difficulty-aware reweighting strategy that combines intrinsic target difficulty with rollout correctness, enabling training to focus on challenging targets. Experiments on SurgSeg-14K show that SurgReasoner outperforms the strongest grounding baseline by large margins, with consistent gains especially on hard samples, suggesting a practical path toward reasoning-driven surgical visual perception.