Teacher-Forced Action Loss Can Select Policies That Never Stop in Vision–Language Navigation
Abstract
A vision–language navigation policy succeeds only by reaching the goal and stopping there. During development, though, checkpoints are usually chosen by a cheaper proxy: teacher-forced cross-entropy between predicted and expert actions, scored one decision at a time. We ran a preregistered three-seed study to find out what that proxy selects. It selected policies that never stop. Executed under the official evaluation on 61 held-out episodes, not one of the selected checkpoints terminates a single episode, and two of them behave identically to a baseline that simply moves forward at every step. The proxy gave no warning, and the reason is arithmetic: it averages over decisions in which the stopping action is rare, so a model can be worse at stopping than at anything else, never predict it once, and still be ranked best. Class weighting changes which action the policy repeats without restoring navigation, and single-seed probes that alter the objective, scale the data, or add metric 3D geometry leave the failure in place. A favorable teacher-forced loss is not evidence of embodied competence. All results are simulated in Habitat-Sim on Matterport3D scans [Chang et al., 2017]; no physical robot is involved.