Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery
Abstract
Low-altitude economy has been experiencing rapid growth in recent years, with significant contributions to the global economy. While common drone tasks such as delivery, inspection, and search-and-rescue typically use Global Navigation Satellite Systems (GNSS) to navigate, there is an increasing need for developing alternative solutions as GNSS signals can be easily jammed, spoofed, or unavailable over a prolonged operational time. As such, cross-view geo-localization (CVGL), which matches an oblique drone view to a geo-referenced satellite tile, has emerged as a potent alternative that lets an autonomous drone localize itself when GNSS fails. Despite strong recent progress, three limitations persist in current CVGL methods: 1) global-descriptor designs compress the patch grid into a single vector without separating what is shared across the view gap (layout) from what is not (texture); 2) altitude-related scale variation is implicitly retained in the learned embedding rather than treated as a nuisance to be marginalized out; and 3) multi-objective training relies on hand-tuned scalars over losses that live on incompatible gradient scales. To address these limitations, we propose SkyPart, a lightweight swappable head for patch-based vision transformers (ViTs) that institutes explicit part grouping over the patch grid. SkyPart has four components grounded in established theory: (i) learnable prototypes that compete for patch tokens via a single-pass cosine assignment; (ii) altitude-conditioned linear modulation applied only during training so that the retrieval embedding is altitude-free at inference; (iii) a graph-attention readout over active prototypes; and (iv) a Kendall uncertainty-weighted multi-objective loss whose stationary points are Pareto-stationary. At 26.95M parameters and 22.14 GFLOPs, SkyPart is the smallest among the top-performing methods in our comparison and sets a new state of the art on SUES-200, University-1652, and DenseUAV datasets under a single-pass, no-re-ranking, no test-time augmentation (TTA) protocol. Furthermore, its accuracy gap to the strongest baseline widens under the ten-condition WeatherPrompt corruption benchmark