Decoupled Prototype Contrastive Alignment Hashing for Cross-Modal Retrieval
Abstract
Cross-modal hashing is essential for efffcient large-scale re-trieval in multimedia analysis. However, existing unsupervised methods often struggle to fully capture interactions between modalities, enforce multi-level alignment, and maintain stable optimization. To overcome these challenges, we propose Decoupled Prototype Contrastive Align-ment Hashing (DPCAH), which employs a cross-modal semantic con-trastive learning module that concatenates image and text features and encodes their interactions via a transformer to produce uniffed represen-tations. feature-level and hash-level contrastive objectives jointly align semantic information and guide discriminative hash code learning across modalities. A decoupled prototype consistency module further models cross-modal correlations independently, enhancing semantic alignment while ensuring stable and robust optimization. Experiments on three benchmark datasets demonstrate that DPCAH outperforms state-of-the-art methods in cross-modal retrieval.