LiFT: Likelihood-Free Tree-Structured Policy Optimization for Flow-Based VLAs
Abstract
Flow-based vision-language-action (VLA) models provide expressive action generation for embodied control, but supervised fine-tuning (SFT) often confines them to narrow expert behaviors. Online reinforcement learning (RL) can improve beyond demonstrations through environment interaction, yet flow-based VLAs struggle with sparse long-horizon rewards and intractable action likelihoods. We propose \textbf{\emph{LiFT}}, a likelihood-free tree policy optimization framework for flow-based VLAs. LiFT first addresses sparse-reward credit assignment by creating sibling continuations from shared rollout histories and using their final outcomes to compute branch level relative advantages. LiFT further uses local action-flow instability as an expansion criterion, focusing rollout budget on ambiguous action-generation regions that are most likely to reveal meaningful outcome differences. Finally, LiFT applies branch-level advantages through a decision-level surrogate ratio on executed action chunks, matching policy updates to the granularity of tree-based credit assignment. This yields a likelihood-free policy update without attributing chunk-level credit to individual denoising steps. Evaluations on in-distribution and out-of-distribution benchmarks show that LiFT improves task success and robustness to distribution shifts.