Learning Gaussian Embeddings from Temporal Views of Satellite Image Time Series
Abstract
The analysis of Satellite Image Time Series (SITS) is crucial for understanding dynamic Earth processes, yet effectively pre-training Geospatial Foundation Models (GFMs) on these 4D data remains a challenge. Although self-supervised learning via Joint-Embedding Predictive Architectures (JEPAs) has shown promise in computer vision, their application to SITS oversights the temporal dimension during the learning process. We propose GETSITS, a framework that enforces extraction of Gaussian Embeddings from Temporal views of SITS, making it suitable for a wide variety of Remote Sensing tasks. We pre-trained ViT-T, ViT-S and ConvNeXtV2 architectures on Sentinel-2 images from SSL4EO, and performed a comprehensive evaluation across several benchmarks. Results show that GETSITS learns robust representations compared to other GFMs and architectures pre-trained on the same dataset, on tasks such as crop type mapping, flood segmentation, among others.