FlashSSM-3D: A Distilled State Space Model for Lightning-Fast Dense 3D Reconstruction
Nyle Siddiqui ⋅ Mubarak Shah
Abstract
3D geometric understanding is undergoing a major transition where rigid, deterministic, optimization-based pipelines are being replaced by large-scale, deep learning models. While recent methods achieve impressive performance on tasks such as multi-view 3D reconstruction and camera pose estimation, they suffer from two critical limitations: $\textbf{(i):}$ current state-of-the-art (SOTA) models require industry-level compute and training corpora, severely limiting accessibility, and $\textbf{(ii):}$ recent approaches remain constrained to modifying existing pretrained transformer models, since substantially changing their architecture would require retraining from scratch. We instead study a complementary direction: training an attention-free state space model for dense 3D understanding. We introduce $\textbf{FlashSSM-3D}$, the first 3D reconstruction state space model capable of performing accurate geometric reasoning over hundreds of images with linear computational complexity. However, training such a model from scratch with competitive performance is impractical under typical academic compute constraints. To address this, we train FlashSSM-3D by distilling knowledge from VGGT, avoiding the need to learn 3D geometric priors entirely from scratch. However, standard distillation losses are poorly suited for dense 3D reconstruction, where preserving real-world spatial structure is essential. We therefore propose a $\textbf{geometry-aware knowledge distillation}$ method specifically tailored for large-scale 3D reconstruction models. Our method introduces two novel losses: a Procrustes-based alignment loss that enforces global structural consistency between student and teacher 3D point predictions, and a camera reprojection loss that directly supervises view-consistent camera geometry between teacher and student. Extensive experiments across multi-view 3D reconstruction and camera pose estimation benchmarks demonstrate that FlashSSM-3D achieves competitive performance while requiring only $\sim25\\%$ of the training data and less than $10\\%$ of the training compute used by VGGT. At inference time, FlashSSM-3D benefits from the practical efficiency of state space models, reconstructing scenes from 500+ views in under 10 seconds while maintaining competitive accuracy.
Chat is not available.
Successful Page Load