Megadose Built for builders and researchers.

Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation

· HF Daily Papers ·
ViGAR splits robot manipulation into visual subgoals first, then action generation.

The paper proposes a hierarchical world-action model for long-horizon tasks made of coordinated subtasks. Its planner predicts the next visual subgoal from the current observation and instruction, while its executor generates future visual trajectories and actions around that subgoal. The authors report 82.00% success on RoboTwin Clean and 67.02% on Random, beating the strongest baseline by 12.86 percentage points on average. They also say real-world robot tests covered five compositional tasks and two in-context learning tasks. HF Daily Papers' note

score 4

Categories: Research