AdapToPASS: Ambiguity-aware Adaptive Spherical Transformer for Panoramic Semantic Segmentation
Soumyaratna Debnath ⋅ weiming zhang ⋅ Shriram Damodaran ⋅ Dingwen Xiao ⋅ Lin Wang
Abstract
Spherical Transformers have emerged as a promising framework for panoramic semantic segmentation (PASS) by operating directly on spherical geometry and alleviating projection-induced distortions. However, existing architectures rely on assumptions of canonical spherical structure and stable viewpoints, which are frequently violated in real-world $360^\circ$ imagery due to unconstrained camera motion, introducing significant contextual and geometric ambiguity. Consequently, they lack adaptive mechanisms to model such ambiguity, limiting robustness to unseen spherical transformations. In contrast, biological perception is inherently ambiguity-aware: rather than estimating uncertainty probabilistically, it adapts to fluctuations in cue reliability caused by geometric and contextual variations, enabling stable interpretation under complex visual transformations. Motivated by these observations, we first present a systematic analysis of existing PASS architectures under various unseen spherical transformations. Following this, we introduce AdapToPASS, a novel, bio-inspired Spherical Transformer that adaptively models contextual and geometric ambiguities for robust panoramic semantic segmentation. At the core of AdapToPASS are the Adaptive Spherical Attention (AdaSpA) blocks that dynamically modulate attention according to local contextual ambiguity, mimicking the adaptive, context-driven perception of biological vision. To address geometric ambiguities, AdapToPASS employs Bifocal Spherical Representation that reconciles the trade-off between field of view and spatial resolution; along with boundary supervision to emulate the boundary-sensitive nature of biological vision. We evaluate AdapToPASS in both indoor and outdoor semantic segmentation, where it consistently outperforms prior state-of-the-art methods. We further validate AdapToPASS under unseen spherical transformations, where it demonstrates strong robustness and surpasses the next-best method by +13.38% relative improvement in mIoU on Stanford2D3D and +18.77% on WildPASS. Additionally, we introduce a lightweight variant, AdapToPASS-Tiny, with fewer than 2M parameters, which surpasses compact baselines while retaining robustness to spherical transformations.
Chat is not available.
Successful Page Load