LAMP: Language-Modulated Geometric Preservation for Multi-Modal Object Re-Identification
Abstract
Multi-modal object Re-Identification (ReID) aims to achieve robust all-weather perception by leveraging complementary information across heterogeneous modalities. However, existing methods still suffer from limitations in both feature fusion and feature alignment. In terms of feature fusion, most approaches assume equal importance across modalities and do not explicitly account for the quality differences among modalities in dynamic environments. As a result, the fused representations tend to be suboptimal and less discriminative in complex scenarios. In terms of feature alignment, existing methods typically rely on point-wise feature matching while ignoring relational structures among samples, which suppresses discriminative information and introduces cross-modal inconsistencies. To address these limitations, we propose LAMP, a unified framework that rethinks multi-modal interaction through language-modulated control and geometry-consistent alignment. We first adopt a Semantic-Assisted Feature Encoder (SAFE), which incorporates textual semantics into visual features to enhance local discriminative cues and provide more informative representations for subsequent processing. Building upon this, the Language-Modulated Dynamic Routing (LMDR) module generates modality-aware gating signals prior to feature fusion, enabling the adaptive adjustment of modality weights under challenging conditions. Furthermore, to address cross-modal inconsistencies, we introduce a Fused Geometric Prototype Alignment (FGPA) objective based on Fused Gromov-Wasserstein (FGW) optimal transport. Instead of enforcing strict point-wise alignment, FGPA establishes cross-modal consensus at the identity-level relational geometry, thereby preserving discriminative structures while mitigating inconsistencies across modalities. LMDR improves the stability of multi-modal fusion, while FGPA enhances structural consistency in cross-modal alignment. Together, they improve the discriminative capability of the model. Extensive experiments on three multi-modal object ReID benchmarks demonstrate the effectiveness of our method.