HuRo: Robotizing Human Videos for Scalable VLA Pretraining
HuRo turns human-video interaction data into robot-aligned training episodes, then shows VLA pretraining gains on real manipulation tasks.
The paper introduces a pipeline for converting heterogeneous human videos into robot-style observations and action trajectories. It builds a HuRo dataset of about 630,000 robotized episodes and 142 million processed frames from five human-video sources. In four real-world manipulation tasks, scaling this pretraining data raises overall completion from 51.5% to 80.3%, with out-of-distribution completion rising from 34.9% to 72.2%. The authors say visual robotization helps OOD robustness, while end-to-end pretraining with retargeted actions beats visual-only transfer. HF Daily Papers' note
The paper introduces a pipeline for converting heterogeneous human videos into robot-style observations and action trajectories. It builds a HuRo dataset of about 630,000 robotized episodes and 142 million processed frames from five human-video sources. In four real-world manipulation tasks, scaling this pretraining data raises overall completion from 51.5% to 80.3%, with out-of-distribution completion rising from 34.9% to 72.2%. The authors say visual robotization helps OOD robustness, while end-to-end pretraining with retargeted actions beats visual-only transfer. HF Daily Papers' note
score 5