Endowing Your Vision-Language-Action Model with a Predictive Mind
Abstract
Vision-Language-Action (VLA) models transfer semantic priors from large-scale vision-language models to robotic manipulation, but their backbones are typically optimized for static visual understanding rather than action-conditioned dynamics. Consequently, the features used by the policy may lack predictive information about how the scene evolves under the robot's actions. Existing future-prediction methods either reconstruct pixels, overemphasizing appearance details irrelevant to control, or use auxiliary latent predictors that remain weakly coupled with the policy-consumed backbone features. We propose PredMind, a plug-and-play predictive representation learning framework that treats future prediction as a training signal for the VLA backbone itself. Instead of reconstructing future pixels, PredMind aligns intermediate backbone representations with their future counterparts conditioned on executed actions, shifting supervision from low-level appearance to the model's semantic feature space. This directly injects action-conditioned temporal structure into the representation hierarchy used for control while preserving pretrained VLM priors. All auxiliary prediction components are discarded after training, introducing no architectural changes, additional parameters, or inference latency at deployment. Across simulation benchmarks, real-world manipulation tasks, and multiple VLA backbones, PredMind consistently improves learning efficiency and final performance. Especially on LIBERO, it matches or surpasses a fully trained base VLA using only 1/60 of the training iterations, demonstrating the benefit of embedding predictive dynamics directly into backbone features used for control. Anonymous code is available at https://anonymous.4open.science/r/PredMind.