Enabling VLA Action Self-Verification via VLM Token Probability Bucketing
Chen Zhao ⋅ Zhuoran Wang ⋅ Haoyang Li ⋅ Guanlin Li ⋅ Shifeng Bao ⋅ Youhe Feng ⋅ Yang Li ⋅ Jie Tang ⋅ Jing Zhang
Abstract
Test-time sampling can improve Vision-Language-Action (VLA) policies, but only if the model can efficiently select a good action from multiple candidates. Prior work often relies on separately trained external verifiers, adding extra models and inference overhead. We propose \textbf{Token Probability Bucketing (TokenPB)}, a self-verification approach that discretizes continuous action chunks into bucket tokens and co-trains a single VLA checkpoint with a joint objective: a flow-matching loss for action generation and an auxiliary bucket loss $\mathcal{L}_{\text{bucket}}$ that teaches the VLM backbone to model observation-conditioned distributions over bucket-token sequences. At inference, TokenPB samples $M$ candidates, scores each via teacher-forced evaluation under the learned distribution, and selects the best---enabling efficient batched scoring with observation-prefix KV-cache reuse. We provide theory connecting the training objective to a monotonic ranking property and bounds on expected selection regret that separate candidate-set effects from calibration. Across multiple backbones and benchmarks, TokenPB yields consistent success-rate improvements over single-sample inference from the same checkpoint (+2.7 in simulation, +17.1 in real-world, +16.4 under out-of-distribution shifts), and further improves over external-verifier baselines under the same sampling budget.
Chat is not available.
Successful Page Load