Where to Look Matters: Rethinking Sub-Volume Sampling in 3D Medical Self-supervised Learning
Abstract
Self-supervised learning on 3D medical images commonly relies on random sub-volume sampling during pretraining. However, random cropping implicitly treats all spatial regions as equally worth observing, despite anatomical priors and spatial redundancy that make regions inherently unequal in their value for representation learning. To this end, we first show through controlled crop-quality manipulations that better sub-volumes lead to better downstream transfer, revealing the observation policy as an overlooked bottleneck in 3D medical SSL. Motivated by this finding, we revisit sub-volume sampling as a where-to-look problem. We propose VolumeProbe, a plug-and-play framework that replaces random cropping with an offline coarse-to-fine probe bank guided by anatomy-awareness, informativeness, and diversity. We instantiate VolumeProbe as a lightweight Lite variant using geometric and image-statistical cues, and a diagnostic Feat variant using frozen visual features, to examine whether lightweight volume-intrinsic priors suffice or richer semantic cues reliably provide additional value. Extensive experiments across multiple 3D medical SSL backbones and downstream tasks demonstrate that VolumeProbe-Lite consistently improves transfer performance over random sampling. Mechanistic analyses further attribute these gains to more anatomically plausible and informative observations, reduced spatial redundancy, more efficient learning dynamics, and more structured representations.