ProMo3D: Probing Motion Cues in Frozen 3D Foundation Models
Xiaoang Zhang ⋅ Mert Kiray ⋅ Yordanka Velikova ⋅ Manoj Biswanath ⋅ Benjamin Busam
Abstract
Feed-forward 3D foundation models have recently reshaped multi-view reconstruction, yet dynamic 4D understanding still often relies on large-scale motion supervision or hand-crafted, architecture-specific heuristics. We ask whether frozen 3D foundation model tokens already encode motion-relevant information that can be exposed by lightweight readout functions, rather than re-learning motion-specific features with high-capacity decoders. Through probing on three representative backbones, we find that static--dynamic separation is decodable from frozen token representations. Our rank-restricted and spectral analyses on $\pi^3$X reveal that the recovered motion cues are concentrated in a low-dimensional subspace. Building on this, we propose Probing-based Motion Distillation (PMD), which converts pairwise motion-state compatibility scores against a fixed anchor token into dense motion masks. Trained on only 160 YouTube-VOS videos using a single 16 GB GPU, PMD transfers zero-shot to DAVIS, SegTrackv2, and FBMS-59 without target-benchmark fine-tuning. With multi-view pretrained geometric backbones, PMD outperforms training-free heuristics and approaches substantially heavier supervised and self-supervised methods, suggesting that motion decoding from frozen 3D foundation models can be lightweight and data-efficient.
Chat is not available.
Successful Page Load