Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs
Abstract
Vision-Language-Action models (VLAs) have shown strong potential for general-purpose robotic manipulation, yet their closed-loop reliability often degrades under local deployment conditions. Existing evaluations typically treat test episodes as independent zero-shot trials, whereas real robots often operate repeatedly in the same or slowly changing environments, where successful executions provide environment-verified evidence about reliable behavior patterns. We study this persistent-deployment setting and ask whether a partially competent frozen VLA can improve its reliability by reusing its own successful test-time experience. We propose an online success-memory guided test-time adaptation framework for generative VLAs. During deployment, the robot stores progress-calibrated successful observation-action segments in a long-term memory. At inference time, it retrieves state-relevant successful action chunks, filters action-inconsistent candidates through trajectory-level consistency, and aggregates the filtered candidates into an elite action prior. To incorporate this prior into action generation, we introduce confidence-adaptive prior guidance, which injects the elite prior into an intermediate state of the flow-matching action sampler and adjusts the guidance strength according to retrieval confidence. This design allows the frozen VLA to exploit environment-specific successful experience while preserving observation-conditioned generative refinement. This retrieve-then-steer mechanism enables lightweight, non-parametric test-time adaptation without updating model parameters or modifying the generative solver. Experiments in simulation and real-world manipulation demonstrate improved task success and closed-loop stability, particularly in long-horizon and multi-stage tasks.