Beyond Data Scaling: Representation-Centric Pre-training for Vision-Language-Action Models
Abstract
Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are fundamentally harder to scale than web-scale image-text data because they require embodied collection and sparsely cover the physical world. This makes representation quality a central bottleneck: under a fixed robot-data budget, VLA pre-training must convert limited trajectories into transferable visual-action knowledge, rather than merely fit actions. We propose VLAct, a VLA-oriented VLM backbone trained with a representation-centric pre-training recipe. VLAct preserves the broad VLM prior, avoids over-specializing the backbone to a single action head, and encourages shared action semantics across embodiments through VLM-prior preservation, multi-head continuous action co-supervision, and a partially unified cross-embodiment action layout, while leaving downstream users free to attach task-specific action heads during fine-tuning. Across multi-embodiment simulation benchmarks, real-world robot experiments, and unseen-embodiment transfer, VLAct consistently improves downstream performance under fixed fine-tuning protocols. On LIBERO-Plus and RoboTwin 2.0, VLAct surpasses large-scale industrial VLA systems such as ABot-M0 and LingBot-VLA, achieving 82.6\% and 92.5\% success, respectively. Most notably, on RoboCasa-GR1, a humanoid embodiment never seen during pre-training, VLAct with only 20\% of downstream trajectories already outperforms the full-data GR00T-N1.6 baseline. These results are obtained using fully open-source data and only a 16-GPU training setup, showing that representation-centric pre-training is an important independent axis of VLA progress beyond data scaling. All models and training pipelines will be open-sourced.