Never Stop Learning: Test-Time Reinforcement Learning for Vision-Language-Action Models
Abstract
Test-time compute scaling offers a promising path for improving foundation models beyond training-time data and parameters. For Vision-Language-Action (VLA) models, however, applying this paradigm through online reinforcement learning remains challenging due to two coupled limitations. Candidate chunks are often biased toward the pretrained policy's high-likelihood regions, while sparse outcome-level feedback lacks process grounding, making reward signals weakly discriminative and exploration insufficiently guided. To address these challenges, we propose EXPLORE, a test-time reinforcement learning framework that coordinates Exploitation and Exploration for action-chunk generation and drives policy optimization with physics-grounded process rewards. Through this joint design, EXPLORE broadens the effective action space without abandoning the pretrained prior and derives faithful process-level rewards from dense physical interaction feedback to guide online policy improvement. Experiments on SimplerEnv-WidowX, DexMG, and RoboCasa validate the effectiveness of EXPLORE, with consistent improvements across autoregressive and diffusion-based VLA backbones.