TEDI: Unpaired Expert Distillation for Efficient GeoFMs Across Modalities
Abstract
Geospatial foundation models are expensive to pretrain, yet model families often repeat this process across backbone sizes and sensing modalities. We study feature distillation as a cheaper way to reuse existing representations. First, frozen CROMA and TerraMind teachers supervise smaller Vision Transformers through patch-level feature matching. On a nine-task PANGAEA evaluation, the distilled TerraMind Tiny, Small, and Base models outperform released models of the same size by 2.76, 3.02, and 1.87 mean IoU points, respectively; the distilled CROMA-Base gains 2.05 points. Second, we introduce TEDI, an unpaired multi-expert method that combines a multispectral TerraMind expert with a high-resolution RGB DINOv3 expert. A shared Transformer uses modality-specific input embeddings and training-only output heads, allowing the student to learn from independent, non-coregistered corpora. TEDI-Large achieves the highest single-model PANGAEA-9 mean (62.51 mIoU), while TEDI-Base reaches 61.75 with 72% fewer parameters. A wider 90M-parameter Base reaches 62.32, within 0.19 points of the 303M-parameter model. Distillation therefore improves the accuracy-compute tradeoff while combining complementary EO experts.