Concord-SLAM: Cross-Render Concordance for Boundary-Native Semantic Gaussian SLAM
Abstract
Three-dimensional Gaussian Splatting (3DGS) has recently become a promising representation for dense SLAM, due to its explicit scene parameterization, efficient differentiable rendering, and high-quality visual reconstruction. Extending 3DGS to semantic SLAM enables the construction of maps that are not only geometrically accurate but also semantically interpretable. However, existing semantic Gaussian SLAM methods usually optimize semantic, depth, and normal renderings through separate supervision objectives. Such a design underexplores the fact that these outputs are generated from the same Gaussian field and should therefore provide a concordant explanation of the underlying scene structure. This limitation is especially problematic around object boundaries and geometric discontinuities, where inaccurate Gaussian growth can cause cross-boundary expansion, boundary blurring, and semantic leakage. To address this problem, we propose Concord-SLAM, a boundary-native semantic Gaussian SLAM framework that establishes concordance among semantic, depth, and normal renderings while promoting boundary-aware Gaussian growth. Our framework contains two key components. First, Cross-Render Concordance (CRC) jointly examines semantic, depth, and normal views rendered from the same Gaussian map, and converts their structural discrepancies in boundary responses, regional continuity, and local geometry into optimization signals. In this way, disagreement among different renderings is used as an intrinsic cue for improving the semantic-geometric coherence of the Gaussian field. Second, Boundary-Native Splatting (BNS) actively inserts new Gaussians near semantic boundaries, depth discontinuities, and salient geometric variations, encouraging primitives to align with physical and semantic boundaries during map construction rather than correcting boundary mixing only through later optimization. Experiments on Replica and ScanNet demonstrate that Concord-SLAM achieves superior performance in camera tracking, scene reconstruction, and semantic mapping.