Beyond [SEG] Tokens: Training-Free Video Reasoning Segmentation via Counterfactual Inference and Contrastive Concept
Abstract
Video Reasoning Segmentation (VRS) demands both complex semantic reasoning and robust spatiotemporal mask propagation. Existing paradigms typically compress reasoning into an isolated, opaque token (e.g., \texttt{[SEG]}) via expensive fine-tuning, or rely on training-free prompt-then-track pipelines that inevitably suffer from semantic degradation and identity switches when encountering hard distractors. To fundamentally address these vulnerabilities, we propose Counterfactual Logic and Explicit Anchoring Reasoning (\textbf{CLEAR}) framework, a training-free VRS framework inspired by human dual-system cognition. Specifically, we introduce a Causal-Counterfactual Inference mechanism that transforms error-prone coordinate regression into a tractable discrete instance selection task, establishing highly reliable initial anchors via rigorous bidirectional reasoning. During mask propagation, we design an adaptive mode-switching mechanism bridged by a Heterogeneous Event Monitor. The fast-thinking mode utilizes a purely visual engine for efficient temporal propagation, while the Monitor continuously evaluates its reliability across multiple dimensions. Upon detecting propagation anomalies, the slow-thinking mode intervenes and explicitly replaces \texttt{[SEG]} tokens with a structured Contrastive Concept Dictionary for closed-loop identity recovery..Extensive experiments demonstrate that CLEAR significantly outperforms existing training-free methods and achieves performance comparable to state-of-the-art training-based approaches on comprehensive referring and reasoning VOS benchmarks. Code will be available after acceptance.