$\pi^2$: A Simple Framework for 2D-to-3D Registration
Hao Deng ⋅ Xiangtai Yang ⋅ Guangmingzi Yang ⋅ Xijing Wang ⋅ MaYuanxiao ⋅ Sisi Li ⋅ Zhiqiang Tian ⋅ Shaoyi Du
Abstract
In this paper, we present $\pi^2$, a simple, efficient, and accurate image-to-point-cloud (I2P) registration paradigm. Instead of relying on intermediate targets such as 2D--3D correspondences, matching confidence, visibility masks, overlap regions, or pose-related proxies, we reformulate I2P registration as point-wise camera-centered coordinate regression. Given a target image and an unaligned LiDAR point cloud, $\pi^2$ directly predicts where each sampled 3D point should lie in the target camera frame. The predicted camera-centered point set is then aligned with the original point set through closed-form SVD-based rigid alignment, yielding the final relative pose without iterative pose optimization or a learned pose-regression head. To improve robustness under large viewpoint changes and sparse outdoor observations, we further introduce a lightweight camera-ray-based regularization objective that provides visibility-aware supervision during training, without requiring an additional visibility or overlap detection network. Experiments on KITTI and nuScenes show that $\pi^2$outperforms the previous SoTA method ICL by 2.47 and 6.72 percentage points in registration success rate, respectively, while running in real time at about 25 FPS, corresponding to an up to 6$\times$ speedup over ICL. Code and pretrained checkpoints will be made available upon acceptance.
Chat is not available.
Successful Page Load