MCAS: Signal-Processing-Based Multi-View Contrastive Learning for Acoustic Sensing
Abstract
Acoustic sensing is a promising technique for human-computer interaction, but existing learning-based systems still depend heavily on large task-specific labeled datasets. Compared with visual data, acoustic signals are difficult to interpret and annotate, making supervision costly and hard to scale. Meanwhile, existing unsupervised representation learning methods, particularly contrastive learning, are largely designed for natural images and rely on data augmentation to form positive pairs. Such strategies can destroy the temporal-physical structure of acoustic signals, causing shortcut learning rather than meaningful representation learning. To address this issue, we propose MCAS, a signal-processing-based multi-view contrastive learning framework for acoustic sensing. Instead of relying on artificial augmentations, MCAS constructs structure-preserving positive pairs from complementary representations derived from different signal processing pipelines of the same acoustic event. Moreover, MCAS is designed to address the heterogeneity of acoustic representations through progressive cross-view interaction enabled by Hierarchical Shared-Token Fusion and a multi-level pretraining objective with physical consistency regularization. Extensive experiments on two typical acoustic sensing tasks demonstrate the effectiveness of the proposed framework. Under low-label settings, MCAS consistently outperforms both supervised baselines and generic self-supervised baselines, while also showing strong cross-task transferability and robustness to corruptions and noise.