VISTA: Support-Anchored Value Targeting for Fast Flow-Based Vision-Language-Action Policies
Hongjie Cao ⋅ Yuxuan Yang ⋅ Yunpeng Mei ⋅ Peng Cheng ⋅ Chenyu Wang ⋅ Jiamin Wang ⋅ Xiaoyi Fan ⋅ Fang Deng ⋅ Gao Huang ⋅ Jie Chen ⋅ Gang Wang
Abstract
Flow-based vision-language-action (VLA) policies provide a compelling recipe for robot control: they inherit broad behavioral priors from large-scale imitation and generate smooth, temporally extended continuous action chunks. The same properties, however, make offline post-training delicate. Mixed offline robot datasets contain expert behavior as well as partial completions and failures, requiring improvement without drifting far from the data support. An offline critic can in principle provide useful directions, but actor-side critic optimization may suffer from extrapolation and push high-dimensional action chunks off-support. In contrast, advantage-conditioned or reweighting-based methods are more stable, but tend to be conservative and restricted to actions already present in the dataset. We propose VISTA, a post-training framework that turns value estimates into support-anchored distillation targets for pretrained flow policies. VISTA queries a critic only at dataset actions, computes a normalized action-gradient direction, and uses this bounded direction to shift clean-action endpoints and velocity targets. A frozen flow teacher preserves the pretrained action prior, while a behavior-cloning anchor keeps the adapted targets close to the data distribution. The shaped supervision is distilled into a time-indexed action denoising (TAD) student, which predicts time-indexed clean-action estimates with a single network evaluation and obtains the executed action by an analytic final readout. We treat the induced target shift as conservative local guidance rather than a global policy-improvement guarantee. Across five BEHAVIOR-1K simulation tasks and five real-world bimanual manipulation tasks, VISTA achieves the best average success rate in our evaluation against behavior cloning (BC), advantage-conditioned BC, IDQL, and flow Q-learning, while reducing action-expert computation by $15\times$ (and end-to-end policy inference by $1.8\times$) relative to the $20$-step flow teacher.
Chat is not available.
Successful Page Load