Structured Velocity-Field Disagreement for Failure Awareness in Vision-Language-Action Models
GEUN YONG KIM ⋅ Nackwoo Kim ⋅ Wonseok Chae ⋅ HOYOUNG YOO ⋅ SangEun Lee ⋅ Kim H Jin
Abstract
Flow-matching vision-language-action (VLA) models generate continuous action chunks for physical interaction, yet they do not natively provide a reliable confidence measure. We study epistemic uncertainty in flow-matching VLAs using velocity-field disagreement (VFD) between independently fine-tuned SmolVLA action experts, retaining disagreement structure across sampled flow times and repeated policy calls. Across three independent two-member ensembles, cross-flow rules under a fixed configuration maintain 87.1--93.1\% true-positive rates while providing lower false-rejection rates and higher precision than global scalar aggregation, and more consistent failure sensitivity than task-specific conformal aggregation. $K$-consecutive and $K$-total remain close across the $K$ sweep, indicating that repeated disagreement matters more than strict contiguity. Cross-call recurrence further improves selectivity by reducing false rejections and increasing precision while preserving high failure sensitivity, especially for permissive within-call rules.
Chat is not available.
Successful Page Load