What Crosses Time? From Recurrent State to Causal Memory in Vision-Language-Action Models
Hyungjun Nam ⋅ Young Kyun Jang ⋅ Byungchul Kim
Abstract
Vision-language-action (VLA) policies with recurrent memory carry latent states across environment steps, but the presence of past information in these states does not establish which components actually influence later behavior. We study this question in $\mu$VLA, whose recurrent state contains 64 memory tokens with 4096 features each. Recurrent-state geometry identifies a rank-one candidate subspace aligned with the uniform token direction. We hold this candidate fixed and test its behavioral role using causal interventions. Past information remains broadly decodable, yet causal influence on memory-dependent behavior is mostly concentrated in this rank-one subspace. Within it, separately fitted rank-16 feature subspaces preserve memory-dependent behavior across two tasks in which a briefly observed cue must guide a later choice. On a qualitatively different task requiring memory of a continuous object position, a separately fitted rank-8 feature subspace also preserves behavior. Thus, although $\mu$VLA exposes a $64\times4096$ recurrent state, with the token and task-specific feature bases fixed, only 16 dynamic coefficients are sufficient in the recall tasks, and 8 are sufficient in the continuous-position task.
Chat is not available.
Successful Page Load