SeMi-Depth: Semantic-Guided Monocular Metric Depth Estimation Across Diverse Scenes
Abstract
Monocular Depth Estimation (MDE) is a pivotal task in computer vision, aimed at predicting a dense depth map from a single RGB image. Perceiving pixel-level accurate depth from a single image across heterogeneous environments is a capability humans exercise without conscious effort, yet building a single artificial model that generalizes across such settings remains a persistent obstacle in MDE. After all, indoor and outdoor scenes differ systematically in depth range, scene geometry, and illumination, so a model trained on a single domain learns domain-specific priors that may fail to generalize across domains. Existing approaches typically address this through carefully designed network architectures and large training datasets; however, scaling model capacity and data requires substantial computational overhead. Furthermore, existing depth estimation techniques may not fully exploit semantic structure, limiting their ability to resolve geometric ambiguities and preserve accurate depth transitions around object boundaries. In contrast, we propose SeMi-Depth, a unified multi-task framework that integrates semantic structural guidance and explicit domain conditioning to achieve robust metric depth estimation across heterogeneous environments. The core idea is to leverage a shared backbone to extract visual representations while jointly guided by dense semantic and explicit domain conditioning as a complementary source of context for depth estimation. Specifically, our approach uses a cross-attention mechanism to align projected depth features with dense semantic features, allowing the depth representation to exploit scene structure and preserve accurate depth transitions around object boundaries. Concurrently, a dedicated domain head generates explicit domain conditioning to adaptively modulate these semantic-aware depth features via Feature-wise Linear Modulation (FiLM). In the model, we use a Swin Transformer-Tiny backbone to extract multi-scale hierarchical feature maps that are shared across all task branches. Each input frame is processed by the shared encoder, and the resulting representations are directed to the semantic, domain, and depth branches. The semantic and depth features are fused using a cross-attention mechanism, yielding semantic-aware depth features. In parallel, the domain head predicts the scene category and generates a corresponding domain embedding. Since the metric depth exhibits different scale characteristics across these environments, the semantic-aware depth representation is subsequently conditioned on the predicted domain. Specifically, we apply Feature-wise Linear Modulation (FiLM) to modulate the semantic-aware depth features with the domain embedding, producing domain-conditioned depth features for the Dense Prediction Transformer (DPT) decoder for metric depth prediction. We train SeMi-Depth jointly on Cityscapes (Outdoor) and ScanNet (Indoor). On Cityscapes, our model achieves an AbsRel of 0.1238, SqRel of 1.247, RMSE of 2.847, and Accuracy(δ3) of 98%. On ScanNet, SeMi-Depth achieves an AbsRel of 0.1178, SqRel of 0.0456, RMSE of 0.2970, and Accuracy(δ3) of 99%. The domain head achieves 99% classification accuracy. Both qualitative and quantitative evaluations demonstrate that cross-attention-based semantic feature fusion produces sharper depth boundaries and improved spatial consistency compared with independent depth decoding. Furthermore, domain-conditioned FiLM provides an explicit mechanism for adapting depth representations across heterogeneous indoor and outdoor environments.