Interpreting World Action Model Behavior Through Imagined-Future Interventions
Abstract
Understanding agent behavior requires more than measuring whether an agent succeeds: we also need to know what information shaped the action and how it flowed through the run. World action models expose a useful test case because they produce predicted futures alongside robot actions, but co-generation alone does not show whether the visible future explains what the agent does. We use two intervention stages to interpret this behavior. Stage 1 measures how strongly and in what direction specific future content changes behavior, and Stage 2 identifies the internal pathway carrying that effect. Across 22 states and six RoboLab tasks, donor future targets redirect Cosmos 3 actions and executed endpoints almost completely toward the donor rollout. Cosmos Policy shows a smaller first-decision-state effect across ten LIBERO-Long tasks. In Stage 2, restoring recipient K/V at predicted-future token positions removes 83–88% of Cosmos 3 steering. The result is a behavior interpretation rather than a task-success score: it identifies what future content changes the agent’s action, which content is insufficient, and which attention pathway carries most of the effect.