Causal World-State Representations in LeWorldModel
Roy Justus
Abstract
Joint-Embedding Predictive Architectures learn latent-space world models that support planning, but how those latents organize world state is an active target of research. We probe every sublayer write site of the LeWorldModel (LeWM) encoder on Push-T, recovering agent position, block position and block angle from the [CLS] residual stream, and verify each representation causally: donor-free low-rank controllers steer the latent, and edits are assessed under autoregressive rollout through the model's own action-conditioned predictor. The three quantities are encoded in three structurally different ways. Agent position is written by a single attention head whose layer, but not presence, varies across training runs; its mean ablation drops decodability to chance or worse, and a rank-4 steering edit retains 96\% progress after four autoregressive steps. Block position and angle have distinct but overlapping rank-12 subspaces carrying a planar position core and a position-dependent circular angle code; their steered edits retain ${\approx}$80% progress at the same horizon. These findings replicate across the released weights and five independently seeded retrainings, on a 100{,}000-state off-policy dataset that separates true state encoding from on-policy correlation.
Chat is not available.
Successful Page Load