Stable3R: Streaming 3D Reconstruction with Stable Geometric References
Abstract
Streaming 3D reconstruction has become a critical component for real-time applications. While recent advances have adapted large-scale offline feedforward models into streaming architectures via causal attention and KV-caching, they suffer from fundamental instability: the irreversible bias in early frames. This limitation stems from two coupled factors: (1) the existing KV-cache-based causal architecture strictly prohibits access to future information, resulting in information-poor early caches; and (2) current methods typically anchor the pose constraints to the initial frame, amplifying the early bias throughout the sequence. To address these challenges, we propose Stable3R, a holistic framework that establishes stable geometric references by enriching historical caches and anchoring poses to prefixes rather than a single initial frame. Specifically, we first introduce a lookahead-augmented attention mechanism, integrating future observations into past caches without violating causal constraints. Second, we adopt a frame-equivariant architecture with prefix-relative supervision, alleviating the reliance on the biased initial frame. Extensive experiments across multiple benchmarks show that our method consistently boosts reconstruction quality within a strictly causal inference schedule.