Focus on the Moment: Real-Time Semantic Failure Detection in Vision-Language-Action Models
Seul Kim ⋅ Jaein Kim ⋅ Dong-Sig Han ⋅ HEE BIN YOO ⋅ Byoung-Tak Zhang
Abstract
In the physical world, vision-language-action (VLA) agents are powerful, yet prone to semantic failures, wherein the motion remains plausible while the policy manipulates the wrong target. While such failures are recoverable on time, existing detectors require complete trajectories and provide only post hoc assessments, determining whether a failure occurred but not when. To focus on the moment, we propose FailMoment, which detects failures at each frame by utilizing only the recent past. This enables trajectory-level, frame-level, and real-time predictions within a single model under an inference latency budget. As training FailMoment requires frame-level labels, we introduce SemGen, a pipeline that generates semantic failures by relabeling successful demonstrations and pinpointing the failure frame from the gripper signal. It produces $645{,}029$ labeled frames from $3{,}148$ trajectories across $80$ RLBench and LIBERO tasks. On seen tasks, FailMoment remains competitive at the trajectory level, and remarkably, it outperforms all baselines at the frame level, reaching an AUROC of $0.993$ on wrong-position failures, compared with $0.692$ for the best baseline. Sim-to-real experiments on RT-1 further suggest that our frame-level formulation provides a strong foundation for adapting semantic failure detectors to real-world robot data.
Chat is not available.
Successful Page Load