Shortcut Learning in Geospatial Foundation Model Embeddings for Cross-Disaster Damage Assessment
Abstract
Random image-level splits can conceal shortcut learning because the same disaster appears in training and test sets. We audit TerraMind, Prithvi-EO-2.0, DOFA, and DINOv3 embeddings on xView2 and BRIGHT through four diagnostic families: event and acquisition context, temporal views, annotation-derived building count, and pre/post correspondence. Replacing a random split with leave-one-event-out (LOEO) reduces mean balanced accuracy from 0.807 to 0.638 on xView2 and from 0.818 to 0.568 on BRIGHT. An image-free event-prior baseline yields IID AUC 0.821 and 0.715. On xView2, balanced accuracy rises from 0.577 with pre-only embeddings to 0.617 with post-only embeddings and 0.638 with the complete pair. On BRIGHT, no temporal input exceeds pre-only performance (0.571), exposing the difficulty of its optical--SAR pairing. A one-number building-count baseline reaches 0.629 and 0.653. Mismatching locations within the same event and class reduces balanced accuracy only to 0.617 and 0.552. Source-only PCA, a supervised MLP, supervised contrastive learning (SupCon), and cross-event SupCon partially recover held-out-event performance: SupCon performs best on xView2 and PCA on BRIGHT. The results support event-held-out evaluation and explicit shortcut controls before interpreting benchmark performance as cross-disaster transfer.