Transferability of Geospatial Foundation Model Embeddings for Crop-Type Classification
Abstract
Geospatial foundation models (GFMs) are proposed as a way to reduce the annotation burden of training deep neural networks for downstream applications. In agriculture, and crop-type identification in particular, the representations produced by GFMs, known as embeddings, are expected to capture crop-related information on their own. While the quality of these embeddings is typically assessed through probing — training a shallow classifier on top of them — labelled data to perform this probing in the region and season of interest is not always available. This raises the question of whether embeddings capture knowledge that transfers across regions and years, rather than only within the domain in which the probe was trained. We test this using frozen TESSERA embeddings across nine European countries for 2018 and across Austria for 2016–2023, comparing lightweight classifiers trained on the embeddings against classifiers trained on the raw satellite time series from which they are derived. Cross-country transfer retains 66\% of in-distribution accuracy for the embeddings, against 59% for the raw series, while cross-year transfer retains 92–94%, against 88%. Embedding-based models also match or exceed the raw-data baselines while using only one quarter of the labels. Probing further shows that country and coordinates, but not acquisition year, are recoverable from the embeddings; domain-adversarial adaptation narrows this gap, improving mean target accuracy by 2.8 percentage points. These results indicate that TESSERA embeddings transfer well across years but, as expected, only partially across countries, confirming that pre-training entangles crop semantics with geographic context.