ForesightFlow: Self-Guided Flow Matching for Improving Vision-Language-Action Models
Abstract
Large vision-language-action (VLA) policies are increasingly trained as conditional generative models over action chunks, yet they typically improve by imitating the data they are given. This is limiting in deployment, where robots collect mixed-quality experience containing demonstrations, partial completions, recoverable mistakes, and failures. Full behavior cloning imitates failures, filtered behavior cloning discards useful sub-trajectories, and offline RL usually requires a large separate critic. We introduce ForesightFlow, a self-guided flow policy that augments each generated action chunk with a learned success-potential vector. The same flow therefore proposes candidate actions and scores them, enabling best-of-K inference without an external critic. The key challenge is that policy improvement and value calibration require different supervision: advantage weighting should suppress low-quality actions, but applying the same weights to the potential dimension removes the failure gradients needed for calibration. We address this with decoupled advantage-weighted flow matching, applying exponentiated advantage weights only to action velocities while training potential velocities uniformly. We also derive a single-step boundary estimator for conditional flow matching with independent endpoint sampling, allowing advantage computation with one stop-gradient forward pass. Across five BEHAVIOR-1K simulation tasks and five real-world bimanual tasks, ForesightFlow improves over imitation baselines, matches the strongest separate-critic baseline in average simulation success, improves real-world success, and reduces training compute by 38\%. Ablations show that decoupling prevents value hallucination, the one-step estimator preserves candidate-ranking fidelity, and self-guided sampling improves long-horizon performance.