HIST3R: Rectified State Decomposition from History for Streaming 3D Reconstruction
Abstract
Streaming 3D reconstruction enables long-sequence scene understanding by processing frames sequentially, but remains challenging when historical information is compressed into a fixed-size recurrent state. Existing training-free extensions of implicit-memory models mainly improve stability by modulating update magnitude, thereby controlling how strongly each incoming observation modifies the state. However, the geometric intent of the update, namely the direction in which the observation drives the recurrent state, remains tied to the current-frame candidate and can accumulate directional bias over long sequences. We propose HIST3R, a training-free framework that improves recurrent state updating through a direction--magnitude decomposition. HIST3R uses historical recurrent states to rectify the update direction, providing a more reliable geometric intent, and uses observation--history visual discrepancy to modulate the update magnitude, yielding an adaptive update gain. This design makes recurrent updates less dependent on isolated frame-level evidence while preserving the bounded-memory advantage of implicit streaming reconstruction. Across camera pose estimation, 3D reconstruction, and video depth estimation benchmarks, HIST3R consistently improves long-horizon performance without additional training. At 1000 input frames, HIST3R achieves relative ATE reductions of 23.87% on ScanNet and 33.65% on TUM-Dynamics over the strongest baseline; at 500 frames, it improves long-sequence 3D reconstruction over the strongest CUT3R-based variant by 56.52% on 7-Scenes and 33.96% on NRGBD.