Megadose AI progress, ranked and analyzed.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

· HF Daily Papers ·
Xiaomi says its VLA robot model scales with more real-world trajectory data and larger model size.

The paper introduces Xiaomi-Robotics-1, trained on more than 100,000 hours of real-world manipulation trajectories. It uses an auto-labeling pipeline to describe scene state changes in natural language, then post-trains the model for robot embodiments and imperative instructions. The authors report stronger out-of-the-box performance in unseen environments after larger-scale pre-training. They also claim new state-of-the-art results on RoboCasa365 and RoboDojo. HF Daily Papers' note

score 7

Categories: Model Releases, Research