GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch
GigaWorld reports 85 ms action-only inference on a local RTX 4090 by avoiding future-video generation at runtime.
The paper frames GigaWorld-Policy-0.5 as an action-centered World Action Model for robot control. It uses future visual dynamics during training, then decodes only actions during inference to cut compute. A Mixture-of-Transformers splits visual dynamics modeling from action generation, so less of the model is active at runtime. The team also says an agent-based AutoResearch pipeline helped search training configurations with less manual tuning. HF Daily Papers' note
The paper frames GigaWorld-Policy-0.5 as an action-centered World Action Model for robot control. It uses future visual dynamics during training, then decodes only actions during inference to cut compute. A Mixture-of-Transformers splits visual dynamics modeling from action generation, so less of the model is active at runtime. The team also says an agent-based AutoResearch pipeline helped search training configurations with less manual tuning. HF Daily Papers' note
score 5