MvFFN: Multi-view Floor-Plan Feed-Forward Network for Unposed Wide-Baseline Panorama Layout Reconstruction
Abstract
Reconstructing accurate camera poses and floor-plan layouts from sparse, unposed, wide-baseline RGB panoramas remains a challenging open problem. Direct multi-room floor-plan prediction is especially difficult in this setting: pose, geometry, semantics, and global layout structure are tightly coupled, yet a complete floor-plan is hard to learn as a single global output. Our key insight is to keep the model's predictions strictly local. Multi-view Ordered Wall Instance Segmentation (MOIS) is designed around this concept. Each panorama column predicts only the wall it observes plus the next two adjacent walls in clockwise order. The global wall partition, room layouts, and shared coordinates emerge from deterministic aggregation of overlapping local predictions. Building on MOIS, we propose Multi-view floor-plan Feed-Forward Network (MvFFN), which jointly predicts coarse camera poses, dense global geometry, and per-wall semantic outputs. A downstream Layout Chaining (LC) post processing pipeline then assembles these predictions into multi-room floor-plans with shared global coordinates, consistent wall identities, and connected room structure. The dense MOIS predictions recover cross-view wall identity and local within-room connectivity more accurately than composed modular baselines, and the assembled layouts substantially outperform prior methods on wall and junction level accuracy. We also introduce the ZInD-CrossView dataset, which augments ZInD with globally unique wall instances and cross-view segmentation labels for this task.