Complexity-Aware LoRA Aggregation for Modality-Heterogeneous Federated Person Re-identification
Abstract
A unified cross-modal person re-identification targets robust cross-modal retrieval with queries from diverse modalities like visible, infrared, and text.Such framework requires multi-modal data and centralized training, raising data costs and privacy risks. Hence, we propose a modality-heterogeneous federated ReID task to learn a unified cross-modal framework from clients with different cross-modal data.However, it faces two challenges.First, varying inter-modality complexity across tasks leads to different retrieval patterns and parameter space. Second, different intra-modality complexity results in imbalanced contributions when merging LoRA modules.Existing federated learning methods use uniform merging and transmit full models, which hinders cross-modal knowledge integration and increases communication overhead. To address these issues, we adopt Low-Rank Adaptation (LoRA), a lightweight module for fine-tuning large models through low-rank updates, enabling clients to send compact parameters.On top of LoRA, we develop MIRAGE (Modality-heterogeneous Intelligent federated ReID Aggregation with Global Efficiency), adapting aggregation to inter- and intra-modality complexity.When merging visible LoRA, MIRAGE selects key ranks to protect retrieval patterns while retaining shared ranks for common ability.As for infrared and text LoRA, it reduces conflicts with an adaptive pruning method. In both strategies, task complexity is incorporated into aggregation to balance contributions. Experiments on cross-modal retrieval tasks demonstrate the superiority of MIRAGE.The code and dataset will be publicly released.