Channel Mixer: A Pretrainable Tokenizer for Scalable Multi-Channel Vision Transformers
Lukas Miklautz ⋅ Lucas Miranda ⋅ Nikola Bulat ⋅ Armin Lambacher ⋅ Dexiong Chen ⋅ Karsten Borgwardt
Abstract
Channel-adaptive Vision Transformers (ViTs) struggle to scale to non-RGB images with many uncorrelated or weakly correlated channels. Existing approaches roll out each channel into its own patch embedding, so the token sequence length grows linearly with the number of channels and self-attention cost grows quadratically. We introduce \emph{ChannelMixer}, a multi-channel tokenizer that aggregates all channels within each spatial patch into a single channel-mixing token, decoupling the token sequence length from the number of channels. ChannelMixer is pretrained once with a lightweight autoencoder in a channel-agnostic manner and the resulting tokenizer can be easily integrated with ViTs trained under both supervised and self-supervised objectives. On multi-channel imaging benchmarks spanning 3 to 29 channels, ChannelMixer matches or outperforms channel-adaptive state-of-the-art methods while using approximately $20\times$ fewer FLOPs and $29\times$ fewer tokens than rolled-out tokenizers, measured on 29-channel images with a ViT-S/16.
Chat is not available.
Successful Page Load