Predictable Before Action, Resistant to Correction: Failure Signals in Vision-Language-Action Models
Abstract
Long-horizon manipulation with vision-language-action (VLA) policies remains unreliable under distribution shift. We ask whether a policy's memory state signals impending failure early enough to act on. On MemoryVLA, a linear probe on the fused cognitive state predicts eventual failure at 0.862 AUC before any episode could have ended, generalizing to unseen perturbation dimensions (0.847) and unseen scenes (0.811); most of this signal is present in the first observation. We test four interventions: Best-of-N sampling, a scripted recovery macro, slowed execution, and memory bank clearing. Each measurably alters the robot's behavior, yet none reliably improves success on this policy and benchmark; an outcome-informed oracle shows that headroom exists. Aborting predicted failures early saves 27% of compute at 2% cost in successes.