Detection Is Free, Diagnosis Is Not: Externally Validated Failure Characterization for Latent World Models
Abstract
A latent world model promises abstraction: an encoder compresses pixels into a state-like code, a predictor advances that code, and what is learned about the dynamics should transfer across changes in appearance. We study the failure of this promise in the encoder and predictor of a JEPA world model, on 25 evaluation cells (combinations of world and adaptation regime) spanning three controlled shifts of a pixel-based pushing task, each cell evaluated in closed-loop planning. Standard self-supervised metrics detect the failures but cannot explain them: they are computed on the distribution the model was validated on, while planning consumes the predictor's own iterates. Two of our models have indistinguishable one-step next-latent prediction error (latent MSE 0.516 vs.\ 0.517) and both fail at closed-loop planning (success rate 0.01); rolling the same metric out five steps separates them by a factor of 1.24 without explaining either. We present an externally validated diagnosis protocol, frozen linear probes decoded over free rollouts with built-in controls, that identifies the failing module (predictor or encoder), the affected state coordinates, the horizon, and the failure mode, and whose diagnoses hold under intervention. We show that fine-tuning only the predictor recovers 81--91\% of the full model fine-tuning gain in planning success when the diagnosis identifies the predictor, and at most 18\% when it identifies the encoder. Finally, we turn the protocol on itself: within any single world, standard label-free metrics rank the same failures just as well. What they cannot do is name the module, and their ranking breaks down when models from different worlds are compared jointly (pooled Spearman 0.58, below a world-identity baseline of 0.75, against 0.85 for the probe).