Long-term Embodied Visual Tracking with Lightweight Vision-Language-Action Models
Abstract
Embodied Visual Tracking (EVT) aims to autonomously control an agent to follow a target in 3D space, which is critical for robotic navigation and human-robot interaction. However, robust long-term EVT in real-world scenarios faces two key bottlenecks: frequent tracking failures caused by occlusions and distractors, and the lack of a long-term benchmark for model training and evaluation. To address these, we propose a systematic solution. First, we develop LT-VLA, a lightweight VLA model featuring a Target Memory module to maintain consistent target identification and an Adaptive Execution module to adjust tracking actions based on observation reliability. Second, we construct LTEVT, a unified long-term EVT benchmark comprising 11 high-fidelity indoor and outdoor environments, enriched with over 5,000 assets and 325 human targets to simulate occlusions and distractors. Along with the benchmark, we release the LTEVT-4500K dataset for large-scale training and a comprehensive evaluation toolkit. Extensive experiments demonstrate that our 0.6B model achieves state-of-the-art performance comparable to 7B-scale models while maintaining superior efficiency. Notably, LT-VLA achieves 62.4% average Success Rate on EVT-Bench and runs at 20 FPS on a Unitree G1 robot, delivering a practical and robust solution for real-world deployment.