CMI-Trans: Cross Modal Inconsistency-aware Transport for HSI-LiDAR Classification
Yanli Li ⋅ Xuan Tan ⋅ Ding Qi ⋅ XINYANG JIANG
Abstract
Joint classification of hyperspectral imagery (HSI) and Light Detection and Ranging (LiDAR) data benefits from complementary spectral--geometric cues, but existing fusion methods largely assume locally consistent cross-modal features and symmetric pixel-wise trust, which is often violated in real urban scenes. We term this pixel-level phenomenon \emph{cross-modal local inconsistency} (CMI) and quantify it with a Strong-Edge CMI Ratio: dual-modal gradient analysis shows that over $16\%$ of strong-edge pixels exhibit single-modality dominance across three benchmarks, while ROI-level analysis reveals the class-entanglement induced by symmetric fusion. To address this issue, we propose \textbf{CMI-Trans} (\textbf{C}ross-\textbf{M}odal \textbf{I}nconsistency-aware \textbf{Trans}port), which promotes per-pixel reliability modeling from a post-hoc diagnostic to a structural signal for alignment and training. CMI-Trans combines Energy-Calibrated Reliability (ECR), Uncertainty-Shaped Transport Alignment (USTA) with Sinkhorn-regularized optimal transport, and Triadic Co-Optimization (TCO) that jointly optimizes classification, uncertainty consistency, and alignment stability on a Mamba-based fusion backbone. On Houston, Trento, and MUUFL, CMI-Trans achieves competitive overall accuracy and substantial gains on CMI-affected classes (e.g., $+17.90\%$ on Houston C10). The ECR-derived uncertainty score also separates reliable from erroneous predictions ($\rho(U,e)<0$; top-$20\%$ highest-$U$ error rate $\le 0.18\%$), while improving feature separability in shadow-induced CMI regions. These results show that explicitly modeling CMI through uncertainty-shaped adaptive alignment yields more accurate and interpretable multimodal remote sensing classification.
Chat is not available.
Successful Page Load