ContinuLoc: Continuous Pose Inference over Neural Fields for UAV Geo-Localization
Abstract
Cross-view UAV geo-localization involves three design choices: how to represent the UAV observation, how to represent the geo-referenced scene, and how to infer pose from their interaction. The dominant paradigm organizes the world as a discretized 2DoF candidate library, coupling scene representation to inference and bounding pose accuracy by sampling resolution. We decouple scene representation from pose inference by encoding the geo-referenced world as a continuously queryable semantic neural field, enabling localization to be formulated as explicit optimization over a continuous 4DoF pose space — parameterized by 2D ground position, in-plane heading, and viewing scale. Pose hypotheses are evaluated globally in a shared semantic feature space, bypassing the sampling bottleneck of retrieval. We learn 4DoF pose-sensitive observation representations under full 4DoF supervision provided by ContinuScenes, a multi-city benchmark with dense near-nadir UAV imagery and complete pose annotations. Inference proceeds hierarchically, progressing from broad pose-space coverage to adaptive concentration on high-compatibility regions. Preliminary experiments demonstrate that our method substantially outperforms prevailing localization paradigms, including discrete retrieval and absolute pose regression, across increasingly strict 4DoF pose criteria.