Perceptive Humanoid Locomotion on Challenging Terrain at Low Speed: A Policy-Shared World Model Predicts Falls but Not Distribution Shift
Nguyen T Huynh
Abstract
Latent world models are commonly trained alongside control policies, but are seldom evaluated by whether a controller can act on their predictions. We study this question in perceptive humanoid locomotion over challenging terrain at low speed. A latent world model is trained as an auxiliary objective on the policy's own depth encoder, and a scalar value head reads its latent state to estimate the probability that the robot falls within the next two seconds; supervision is free, since the simulator already distinguishes failure terminations from time-outs. We report three findings. First, sharing the encoder between policy and world model does not degrade control: held-out success changes by $-0.2$ and $+3.0$ percentage points on two humanoid platforms, and the latent remains non-degenerate. Second, the value head anticipates falls well: AUROC $0.82$ and $0.74$, average precision near $8\times$ the base rate, and a median lead time of $1.5$~s---an order of magnitude earlier than proprioceptive fall detectors. Third, the model's own prediction error does \emph{not} detect distribution shift. On a course containing every terrain the training mix omits, it separates unfamiliar from familiar ground at AUROC $0.513$, because error varies more across terrain types than between seen and unseen ground. The signal is therefore informative about \emph{when} the robot will fall, but not about \emph{where} the model is out of its depth.
Chat is not available.
Successful Page Load