Attribution-Guided Shared-Private Decoupling for Noise-Reduced Audio-Visual Representation Learning
Abstract
Audio-visual representation learning commonly aligns paired audio and visual signals with sample-level contrastive objectives. However, real-world audio-visual pairs usually contain heterogeneous patch tokens: only a subset carries reliable cross-modal shared semantics, while many others mainly encode modality-private residuals. Aggregating all tokens into a single global representation can therefore introduce semantic noise into cross-modal alignment. Motivated by the minimal-sufficiency principle, we propose an attribution-guided shared-private decoupling framework to mitigate this issue. Our key idea is to identify shared-cover subviews that preserve dominant cross-modal information while suppressing modality-private residuals, and use them to supervise decoupled shared and private representations. To identify such subviews efficiently, we draw on transformer attribution method and introduce Layer-Truncated Attribution (LTA), which provides a lightweight estimate of token-level relevance to cross-modal shared information. We further adopt a teacher-student framework, where the teacher provides attribution-induced supervision and the student internalizes shared-private decoupling into its representations, without introducing extra inference cost. Experiments on retrieval, classification, and sound-prompted semantic segmentation show consistent gains over strong baselines. Compared with baseline, our method improves the average zero-shot R@1 by 4.0 across AudioSet and VGGSound retrieval, and improves AS20K classification mAP by 3.1, demonstrating the effectiveness of fine grained attribution guidance for reducing semantic noise.