STEMFly: Enhancing UAV Vision-Language Navigation via Sensor Grounding, Temporal Diversity and Episodic Memory
Guangdao Zhu ⋅ Xu Chen ⋅ Shuhong Hou ⋅ Weili Guan ⋅ Bin Chen ⋅ Yaowei Wang ⋅ Xiang Deng
Abstract
Despite rapid progress in indoor vision-and-language navigation (VLN), its aerial counterpart remains significantly more challenging and underexplored. In this setting, unmanned aerial vehicles (UAVs) must interpret free-form instructions and traverse kilometer-scale outdoor environments. Although recent UAV-VLN methods have shown promising results, they often fail to fully exploit contextual signals across modality, time, and prior experience. This limitation appears in three aspects: the predominant reliance on RGB-only inputs, the use of discrete frame inputs that fail to capture temporal dynamics, and the absence of episodic memory, resulting in unreliable termination decisions. To this end, we propose $\textbf{STEMFly}$, a unified framework grounded in $\textbf{S}$ensor signals, $\textbf{T}$emporal diversity, and $\textbf{E}$pisodic $\textbf{M}$emory. STEMFly comprises three components. First, Sensor-Augmented Prompt Injection (SAPI) incorporates multi-modal sensor signals as semantically grounded language prompts that can be interpreted by the language model. Second, Temporal-Diversity Frame Selection (TDFS) constructs a temporally informative observation sequence via entropy-guided frame sampling. Third, Memory-Augmented Success Verification (MASV) improves termination reliability by validating navigation outcomes through episodic memory retrieval. In addition to methodological contributions, we identify and rectify systematic annotation inconsistencies in the OpenUAV benchmark, releasing a revised evaluation protocol. On the original OpenUAV benchmark, STEMFly outperforms the prior state-of-the-art TravelUAV, achieving absolute improvements of 8.26\% in Success Rate (SR) and 6.70\% in Success weighted by Path Length (SPL), while reducing Navigation Error (NE) by 9.55 meters. Similar consistent gains are also observed on our rectified benchmark. These results demonstrate that holistic contextual grounding, which integrates multi-modal sensing, temporal modeling, and episodic memory, is crucial for improving UAV-VLN performance.
Chat is not available.
Successful Page Load