A Study of Accelerating Flow-Based VLA Inference via a Plug-and-Play Episodic Cache
Ryuji Oi ⋅ Hikari Otsuka ⋅ Kosuke Matsushima ⋅ Yuki Ichikawa ⋅ Masato Motomura ⋅ Tatsuya Kaneko ⋅ Daichi Fujiki
Abstract
Vision-Language-Action (VLA) models have emerged as a promising approach for generalizable robotic manipulations. In particular, flow-matching-based VLA models have shown remarkable success due to their capability to generate precise and smooth action sequences and capture multimodal distributions. However, the iterative denoising process in the action head acts as a major computational bottleneck, posing a critical challenge for real-time deployment. To address this challenge, we propose ActionCache, a plug-and-play external cache that opportunistically reuses past intermediate actions to warm-start generations from the vicinity of target actions, drastically reducing the inference latency. Specifically, ActionCache stores the intermediate actions with compact multimodal keys, which enables retrieval from similar past contexts across different episodes. Experimental results in simulation and real-world environments demonstrate that ActionCache achieves a substantial reduction in the action head latency while maintaining the base model's task success rate in representative flow-matching-based VLAs, $\pi_{0.5}$ and GR00T-N1.6.
Chat is not available.
Successful Page Load