HypMoE-ReID: Hyperspherical Mixture-of-Experts for Large Scale Person Re-Identification
Abstract
With the rapid growth of surveillance data, Person Re-Identification (ReID) is evolving toward large scale scenarios, where datasets exhibit substantial increases in both size and diversity. Under this trend, the widely adopted dense backbone architectures, such as Transformers, suffer from an inherent limitation: Shared parameters couple all input tokens, forcing the model to learn generic representations that are suboptimal for heterogeneous data sources. Although Mixture-of-Experts (MoE) architectures offer a promising solution for handling diverse data, they encounter a fundamental challenge in ReID. Specifically, person images exhibit high inter-instance similarity at the coarse level, while discriminative cues lie in subtle fine-grained details that are easily overlooked. As a result, existing MoE methods often suffer from severe routing collapse, leading to insufficient expert specialization and limited diverse knowledge acquisition capacity. To address these problems, we propose a novel framework, termed Hyperspherical Mixture-of-Experts for Large Scale Person Re-Identification (HypMoE-ReID). Specifically, motivated by theoretical analysis in hyperspherical space, HypMoE-ReID incorporates two key components: (1) a Token Purification mechanism that suppresses noisy tokens detrimental to expert specialization, and (2) a Routing Diversity Learning strategy that explicitly separates experts on a unit hypersphere to enforce specialization and prevent collapse. Furthermore, to facilitate the investigation in large scale ReID, we construct a large-scale ReID benchmark comprising 809,701 images by unifying 13 public datasets with significant distribution gaps. Furthermore, extensive experiments demonstrate that HypMoE-ReID outperforms state-of-the-art MoE-based methods by at least 11.3\%/10.3\% in Average mAP/R@1, and surpasses strong dense backbone models by 2.2\%/2.0\% in Average mAP/R@1. Our code will be released.