RigRecon: Efficient Rig-Aware Street Reconstruction via Dual-Path Spatio-Temporal Interaction
Abstract
Reconstructing multi-camera videos is a fundamental task in autonomous driving. However, recent learned methods either rely on costly attention-based networks for global geometry modeling or require complex post-hoc alignment for chunked predictions, limiting their scalability to long multi-view rig sequences. We propose RigRecon, a generalized rig-aware reconstruction framework that enables efficient cross-view interaction in a low-resolution space while asynchronously decoding high-resolution depth maps in a frame-wise manner. Input images are first down-sampled and processed by a transformer backbone to capture coarse global structure, which is then provided to a decoder to reconstruct frame-wise high-resolution depth maps. By fully leveraging calibration and low-resolution feature interaction, RigRecon achieves consistent-scale geometry and accurate pose estimation with significantly reduced memory cost. It achieves state-of-the-art performance on 3D reconstruction, depth estimation, and camera pose estimation, and can efficiently process videos with over 1000 images in a single feedforward pass.