FIVE-VLA: Fast and EffectIVE Closed-Loop Autonomous Driving with Recurrent Action Memory
Kemal Oksuz ⋅ Alexandru Buburuzan ⋅ Yuhan Yao ⋅ Puneet Dokania
Abstract
State-of-the-art vision-language-action models (VLA) for autonomous driving face critical limitations: excessive parameter counts, inefficient high-resolution image processing, and lack of temporal memory. We introduce **FIVE-VLA** (Fast and EffectIVE VLA) to address these through two key contributions. First, we employ an efficient vision encoder that processes high-resolution ($448 \times 896$) images while generating only 98 tokens, over $5\times$ fewer than existing approaches, and bypass text generation entirely for single-pass trajectory prediction. Second, we propose **Recurrent Action Memory (RAM)**, a lightweight module that conditions action prediction on previous action tokens, providing temporal context critical for manoeuvres such as overtaking and emergency braking. With only 641M parameters, FIVE-VLA completes *$\sim$10\% more routes without traffic rule infractions* than the previous state-of-the-art VLA on the challenging \textit{Bench2Drive} closed-loop driving benchmark. Furthermore, FIVE-VLA runs at $\sim$30 fps on an A100 and $\sim$4 fps on a T4 GPU (proxy to an edge device), representing an *8--30$\times$ speedup* over previous methods. The code will be made public upon acceptance.
Chat is not available.
Successful Page Load