Beyond Pixel-wise Supervision: Local Structure Regularization for Semantic Segmentation
Abstract
Visual foundation models (VFMs) provide strong region-level representations for semantic segmentation, substantially improving high-level semantic understanding. However, accurate boundary recovery remains challenging, especially for fine structures and semantic transitions. A key reason is that conventional pixel-level supervision largely treats pixels independently and lacks explicit modeling of local structure and neighboring-pixel relationships. To address this issue, we propose a Local Structure (LS) regularization framework that constrains the second-order local organization of semantic transition regions. Unlike existing boundary losses that mainly supervise contour location, distance, overlap, or topology, LS models boundary neighborhoods through a structure-tensor representation of semantic probability fields. It decomposes local boundary structure into two complementary cues: Local Structure Magnitude (LSM), which measures the strength of anisotropic class transitions, and Local Structure Orientation (LSO), which describes their dominant geometric orientation. LS is model-agnostic, compatible with different VFM-backbone-decoder combinations, and introduces no extra inference cost. Extensive experiments on four diverse datasets show that LS consistently improves boundary-sensitive performance while also enhancing overall IoU, demonstrating its effectiveness for boundary-structure learning in semantic segmentation.