SPEXT: A Decoupled Multi-Spectral Foundation Model for Earth Observation and Vision-Language Grounding
Abstract
Vision-Language Models for Earth Observation are bottlenecked at the visual encoder. Existing pipelines either compress multi-spectral signals into RGB pixels or collapse them into a single CLIP-style global vector, losing spectral information or spatial structure respectively. We show this is a consequence of coupling contrastive supervision to the dense feature stream, which degrades spectral fidelity. SPEXT resolves this through architectural decoupling. A wavelength-parameterized encoder produces multi-spectral patch tokens for dense tasks and VLM grounding, while a dedicated alignment token absorbs the contrastive gradient for retrieval. SPEXT simultaneously leads the PANGAEA segmentation benchmark (+2.51 mIoU), tops zero-shot retrieval among multi-spectral CLIP baselines (+4.6 mAP@100), and enables a frozen VLM, paired with a small projection adapter and LoRA, to answer free-form Earth-Observation questions directly from its tokens, substantially outperforming an RGB-pixel baseline on a downstream EO-RAG benchmark.