Not All Features Are Created Equal: A Mechanistic Study of Vision-Language-Action Models
Abstract
A fine-tuned Vision-Language-Action (VLA) policy will pick up the alphabet soup and place it in the basket on demand, then drop the soup off the table when an evaluator shifts the basket five centimeters left. We use activation injection to ask what the policy is actually doing: inject task A's activations into task B's scene at the action expert, and π₀.₅ executes task A's motor trajectory in 99.6% of episodes (n=1,968); X-VLA does so in 99.8%. The injected program is bound to absolute workspace coordinates rather than the visible scene, which mechanistically explains the perturbation brittleness reported in concurrent benchmark work. The same intervention framework applied across six VLAs (π₀.₅, OpenVLA-OFT, X-VLA, SmolVLA, GR00T N1.5, and ACT as a language-free control) on LIBERO, MetaWorld, SimplerEnv, and ALOHA over 420,000+ rollouts surfaces three further findings: language is encoded by every architecture yet behaviorally ignored when vision identifies the goal; SAE pooling preference splits along architecture lines (π₀.₅ per-token, X-VLA mean-pool, SmolVLA indifferent); and pathway specialization replicates wherever expert and VLM pathways are separable. SmolVLA's interleaved fusion attenuates the headline to 52.1% LIBERO override, scoping universality. We will release 424 trained SAEs and an interactive feature-exploration platform on acceptance; an anonymized snapshot is included in the supplementary material.