Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models
VLAct argues that better robot-action representations can beat simply adding more robot data.
The paper presents a continued pre-training approach for vision-language-action models using heterogeneous, multi-embodiment robot data. It reports gains across simulation, real-world tasks, and transfer to unseen embodiments under fixed fine-tuning protocols. VLAct reaches 82.6% on LIBERO-Plus and 92.5% on RoboTwin 2.0, and on RoboCasa-GR1 it beats a full-data GR00T-N1.6 baseline using 20% of downstream trajectories. The authors say the models and training pipelines are publicly available. HF Daily Papers' note
The paper presents a continued pre-training approach for vision-language-action models using heterogeneous, multi-embodiment robot data. It reports gains across simulation, real-world tasks, and transfer to unseen embodiments under fixed fine-tuning protocols. VLAct reaches 82.6% on LIBERO-Plus and 92.5% on RoboTwin 2.0, and on RoboCasa-GR1 it beats a full-data GR00T-N1.6 baseline using 20% of downstream trajectories. The authors say the models and training pipelines are publicly available. HF Daily Papers' note
score 5