Better Read Than Said: Locating Spatial-Planning Failure at the Output Interface of Vision-Language Models
Abstract
Vision-language models (VLMs) are increasingly used as the decision layer of navigation and embodied agents, and they are accurate on many of the perceptual sub-tasks such an agent needs. Open VLMs are much weaker on visual spatial planning, in the parameter range that carries most of the empirical work on it. This paper asks where their failure originates: in the representation the model forms of the physical scene, or in the step that turns such a representation into a specific sequence of actions. We work on grid mazes from the Visual Spatial Planning (VSP) benchmark, where breadth-first search on the ground-truth map gives the set of optimal moves for every state, so every proposed move is scored exactly. We compare five ways of getting a next move out of one frozen hidden state, on the same instances, under the same prompt and from a single forward pass. Two of them are the model's own output: the plan it writes, and the argmax over its four move tokens. Three of them are classifiers we train on the frozen activations to predict an optimal move, with the model emitting nothing. We report optimal-action agreement, the rate at which the chosen move lies on a shortest safe path, and episode success, the rate at which the emitted plan reaches the goal without entering a hole. Across six open models in three families, the classifiers reach optimal-action agreement far above chance and far above what five of the six models write; three of those five return the same first move on every instance we inspected. The written plan improves with parameter count and the classifier does not, so within one family the difference closes from the output side. The difference also reaches the agent's own coordinates, one level below the action: a classifier recovers them from the frozen state, and the models misreport them when asked. It appears in perception as well, where consuming the model's own hazard estimates as a hard constraint scores below ignoring them, and charging the identical estimates as a traversal cost scores above both. Scale buys the ability to report a spatial estimate. The estimate is already present at the smallest size we measured. An evaluation that scores only emitted text understates what a small open VLM represents.