GaitLingo: Self-Supervised Gait Representation Learning with Language Priors
Abstract
Self-supervised gait representation learning aims to learn transferable representations that generalize to unseen scenarios. However, existing methods rely solely on visual self-supervision, making it difficult to assess the semantic value of individual unlabeled sequences and limiting their ability to exploit latent semantic cues. To address this limitation, we introduce language priors into self-supervised gait pretraining. We construct GaitLP-1M, a million-scale silhouette-based gait pretraining dataset without identity labels. A two-stage caption generation pipeline further produces gait descriptions from the corresponding RGB sequences. To the best of our knowledge, GaitLP-1M is the first million-scale silhouette gait dataset paired with gait-centric captions. Building on this foundation, we propose GaitLingo, a language-guided self-supervised framework for gait representation learning. Specifically, we design a Caption-Guided Sample Weighting (CSW) mechanism to prioritize informative samples based on sequence-level semantic richness and attribute rarity. We then introduce a lightweight Temporal Adapter (TA) to better capture motion semantics and improve cross-domain robustness. Finally, we propose Soft Relation Distillation (SRD) to transfer caption-derived relational structures into the visual embedding space. Extensive experiments demonstrate that GaitLingo consistently outperforms prior self-supervised gait pretraining methods in the zero-shot setting across six benchmarks. Compared with GaitSSB, it achieves Rank-1 gains of +12.1% on Gait3D and +7.9% on GREW, as well as an average improvement of +9.1% across four CCPG settings.