Disentangling Channel Semantics in Vision Transformers via Token Decorrelation and Composition-Aware Modulation
Daeun Kim ⋅ Hyejin Park ⋅ Hyesong Choi ⋅ Dongbo Min
Abstract
Multi-channel images encode heterogeneous channel semantics aligned over a shared spatial structure, which requires models to jointly capture globally shared structure and channel-specific variations. Recent multi-channel Vision Transformers attempt to address this challenge by augmenting the `[CLS]` token with memory tokens to organize multi-channel representations under varying channel compositions. However, we find that these context tokens collapse onto a few dominant channels under high inter-channel redundancy, producing redundant representations and a biased global summary that fails to reliably support composition-aware channel interpretation. We propose **MuCa-ViT (Multi-Channel Composition-Aware Vision Transformer)**, a unified framework that addresses these limitations through two coupled mechanisms. **Factorial Token Learning (FTL)** enforces token-level decorrelation among context tokens, encouraging the `[CLS]` token to capture globally shared semantics while memory tokens preserve complementary channel-specific cues. **Channel Composition-Aware Modulation (CAM)** uses the FTL-refined `[CLS]` as a semantic anchor to generate sample-wise modulation signals, which enable composition-aware interpretation of channels within a unified backbone rather than relying on static channel embeddings. Experiments on JUMP-CP, CHAMMI, and So2Sat show that MuCa-ViT outperforms prior multi-channel ViTs across all three benchmarks, and the magnitude of improvement aligns with the inter-channel redundancy of each dataset ($\rho \in [0.19, 0.77]$), validating the diagnostic analysis that motivates our design.
Chat is not available.
Successful Page Load