How Do World Action Models Use Their Imagined Futures? Evidence from Mechanistic Interventions
Abstract
World action models produce predicted futures alongside robot actions, but co-generation alone does not causally establish how the spatial-temporal future affects the action. Prior work shows that behavior depends on information at predicted-future token positions, but not how particular future content shapes the resulting action. We use two intervention stages to characterize embodied spatial reasoning in these models. Stage 1 measures how strongly and in what direction future content changes behavior, and Stage 2 identifies the internal pathway carrying that effect. Stage 1 starts with two normal rollouts from the same simulator state. We insert the future representation associated with one rollout, the donor, into the other, the recipient, hold all other inputs fixed, and measure whether the recomputed action moves toward the donor action, which the model never sees. Across 22 states and six RoboLab tasks, donor future targets redirect Cosmos 3 actions and executed endpoints almost completely toward the donor rollout. Cosmos Policy shows a smaller effect at the first decision state of all ten LIBERO-Long tasks. Matched controls rule out generic perturbation and replacement as complete explanations. Content interventions show that Cosmos Policy’s early-state effect is wrist-camera dominated and consistent with camera-visible robot motion, whereas isolated robot or object pixels do not reproduce Cosmos 3’s full-future effect. In Stage 2, we keep the donor future inserted but restore the recipient’s attention keys and values at predicted-future token positions. This removes 83–88% of Cosmos 3 steering in every tested state, identifying a pathway that carries most of the future content’s causal influence.