Generalised Distillation of Geospatial Foundation Models
Abstract
Geospatial Foundation Models (GeoFMs) are large-scale models with millions of parameters, pre-trained on vast amounts of satellite data to support diverse downstream tasks on earth observation (EO). However, they are typically computationally demanding, which restricts deployment in resource-limited environments. Furthermore, they are predominantly trained using optical data despite the inherent limitations of this modality due to natural phenomena such as cloud cover and atmospheric conditions. Our work aims to leverage multiple modalities by specifically transferring knowledge from a multimodal teacher using optical and radar data, to train a radar-only student model. To ensure computational efficiency, we apply offline feature-based knowledge distillation to a frozen TerraMind-B teacher to compress its multimodal knowledge into a lightweight unimodal TerraMind-S student model. Despite using a smaller student architecture, our distillation approach outperforms TerraMind's state-of-the-art (SOTA) performance on radar data. In particular, we obtain an improvement in mean intersection over union (mIoU) of 11.11\% and 2.29\% on the Crop Type Mapping South Sudan and Sen1Floods11 datasets, respectively. Lastly, we find that the later layers are more important in enhancing the performance of the student model and that distillation performance is thematically dependent.