What, Where, and Boundary: Hierarchical Cognitive Decomposition for Echocardiography Video Segmentation
Abstract
Segmenting the left ventricle in echocardiographic videos remains difficult because the target undergoes continuous deformation and rapid displacement across the cardiac cycle, while low tissue contrast further obscures its endocardial boundaries. Expert echocardiographers navigate these challenges through a hierarchical cognitive workflow, sequentially addressing three fundamental questions:What is the target? Where is it now? Where does the boundary lie? Yet existing methods predict masks end-to-end from temporal features without explicitly modeling this cognitive process. We propose Hierarchical Cognitive Decomposition (HCD), which decomposes the segmentation task into three stages that emulate this expert interpretation workflow. To capture invariant target identity amid changing appearances, Static Identity Anchoring (SIA) anchors the first annotated frame as a semantic reference, enabling stable recognition across the cardiac cycle. To maintain spatial focus despite inter-frame motion, Adaptive Gaze Prior (AGP) converts the previous frame's prediction into a spatial attention prior that localizes the current target position. To resolve boundary ambiguity under low tissue contrast, Structural Boundary Perception (SBP) extracts multi-scale boundary cues within the localized region. On the CAMUS and EchoNet-Dynamic benchmarks, HCD achieves state-of-the-art performance with only 1.35M trainable parameters (out of 35.2M total) at 68 FPS, demonstrating that explicitly encoding clinical cognitive priors into network design yields both effective and efficient segmentation.