A Safe Action Is Not Enough: Feasible-Future Decoding for Vision-Language-Action Policies
Tu Nguyen ⋅ Matthieu Zimmer ⋅ Vu A Vu ⋅ Xuebing Zhou ⋅ Haitham Bou Ammar
Abstract
A locally safe action can still be the first step into failure. A candidate may be likely under a frozen vision-language-action policy and satisfy every available local check, yet leave no policy-supported route to safe task completion. We call this the **feasibility--likelihood gap**. Starting from the history-conditioned trajectory law induced jointly by the policy and environment, we derive the exact next-block marginal of a prior-preserving distribution restricted to safe completion. The marginal exposes a **feasible-future mass** that answers two questions: does any policy-supported safe completion remain after the action, and how much weighted continuation mass remains if it does? Because this quantity is generally unavailable at decision time, we study a selective finite-candidate approximation. Our analysis treats the decision to invoke reranking separately from the approximation used after intervention, and gives conditions for recovering the best viable candidate that remains available. We instantiate this principle as an alarm-triggered, training-free candidate reranker. On Safety-CHORES, the generic configuration lowers mean cumulative safety cost by $1.9$%–$57.5$% across six settings while remaining within $2.5$ percentage points of policy-sampling success and $0.82$ mean steps of exposure. Together, the analysis and results connect an exact policy-relative safe-completion target to a practical selective decoder that improves the evaluated operating points without retraining the policy.
Chat is not available.
Successful Page Load