Learning Active Perception and Manipulation via Spatio-temporal Visual Memory
Enshen Zhou ⋅ Mengzhen Liu ⋅ Yibo Li ⋅ Yanjun Ding ⋅ Yuheng Ji ⋅ Pengwei Wang ⋅ Zhongyuan Wang ⋅ Lu Sheng ⋅ Shanghang Zhang
Abstract
Active perception is crucial for robots to interact with unstructured scenes, where a closed perception-action loop is required. In this work, we introduce ActiveZero, an end-to-end framework that formulates active perception as information-driven spatio-temporal memory management. Leveraging memory as a universal information interface for conveying perception and action, it maintains a visual memory bank, actively expands it via purposeful exploration for missing information, retrieves relevant evidence for execution, and filters memory online for efficiency. This formulation not only enjoys long-context reasoning ability but also enables multi-level supervision for heterogeneous data, going beyond prior paradigms. To support it, we construct ActiveMem, an active perception dataset of 4M episodes with camera actions (300× prior) across 49 scenes for long-horizon tasks (up to 4 subtasks). In addition, we present ActiveBench, the first memory-centric benchmark suite for active perception, covering VQA and simulated manipulation. Experiments show that ActiveZero achieves SOTA on existing active-perception-related benchmarks and ActiveBench. It also outperforms all baselines on challenging real-world tasks by a large margin, even surpassing $\pi_{0.5}$ by 34.16%. Notably, ActiveZero exhibits diverse active perception strategies (e.g., search, track, interact) and handles long-horizon tasks in unstructured scenes with a single unified model.
Chat is not available.
Successful Page Load