GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation
GE-Act 2.0 was trained from scratch on manipulation data, then tested zero-shot across 100 robot tasks.
The model combines a compressed action-aware visual encoder, a one-step future-state planner, and an inverse dynamics module. Scaling co-training data from 300 to 30,000 hours lifted success to 44.1% on G1-OP and 31.1% on G2-90D. The paper says gains appeared across nearly all skill groups, with evidence of cross-embodiment transfer despite limited G2-90D data. It also reports strong grounding of object, color, shape, and position references under the same protocol. HF Daily Papers' note
The model combines a compressed action-aware visual encoder, a one-step future-state planner, and an inverse dynamics module. Scaling co-training data from 300 to 30,000 hours lifted success to 44.1% on G1-OP and 31.1% on G2-90D. The paper says gains appeared across nearly all skill groups, with evidence of cross-embodiment transfer despite limited G2-90D data. It also reports strong grounding of object, color, shape, and position references under the same protocol. HF Daily Papers' note
score 5