ValueFormer: Stage-Aware Value Labels for Online Failure Detection and VLA Post-Training
Inkyu Sa ⋅ Konstantin Stulov ⋅ Rajat Bhageria
Abstract
Generalist vision-language-action (VLA) policies now transfer to new scenes with little task-specific training, but they still fail silently: from the action stream alone, a collapsing rollout looks much like one making clean progress, because imitation supplies no notion of progress. The cheap alternative, a terminal success / failure bit, is learnable in principle yet too sparse to say \emph{when} a rollout went wrong. We argue that the per-frame label is the hard part, and that it works better when continuous and non-monotonic: a score rather than a single success or failure, and one free to fall and then rise again as a rollout goes wrong and recovers. Such a label is cheap to obtain: a person marks only the stage at which an episode went wrong, and the per-frame values are computed from that. We present ValueFormer, a policy-agnostic causal transformer ($3.5$\,M parameters) over a frozen DINOv3 backbone, emitting two per-frame signals in one forward pass: a smooth Monte Carlo value for advantage estimation and a sharp binary value for online mistake detection. Failed episodes carry a stage-aware, success-then-decay return. A teleoperation station already records an intervention flag; reusing it lifts held-out detection from $0.38$ to $0.82$ average precision and flags $95\%$ of mistakes before the operator reacts. On a real bimanual sandwich station, feeding the critic back as a per-frame post-training weight raises task completion from $70\%$ to $85\%$ and removes a repeat-pick failure mode in our trials, although at $n{=}20$ that gain sits within noise.
Chat is not available.
Successful Page Load