SurVAL: Rollout-Free Survival Validation for Robot Policies
Abstract
Reliable evaluation of robot manipulation policies demands costly closed-loop rollouts on hardware. Yet every training run yields dozens of candidate checkpoints, and practitioners can afford to evaluate only a fraction of them. Existing rollout-free proxies for deployment performance, such as average prediction errors, fail to reflect deployment performance because averaging dilutes a single catastrophic deviation by mixing it with many nominal ones, even though a single such deviation can cause a real rollout to fail. To resolve this, we propose Survival VALidation (SurVAL), a rollout-free validation metric that scores each policy by how well it survives along held-out demonstrations. At each step, SurVAL compares the policy's prediction error against a state-conditional threshold automatically derived from demonstration data and converts this into a per-step survival probability. These probabilities are aggregated multiplicatively along the trajectory, so that even a single severe deviation cascades through all subsequent steps, yielding a score that naturally penalizes catastrophic errors without discarding the rest of the trajectory. On diverse tasks in both simulation and real-robot setups, across a diffusion policy and a large-scale vision-language-action model, SurVAL correlates positively with deployment success rate, while baselines degrade or invert the true ordering.