Megadose AI progress, ranked and analyzed.

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

· HF Daily Papers ·
WAA lets VLMs rehearse and revise robot actions in a visual workspace before execution.

The paper describes a multi-agent harness where VLMs control robot manipulation through contact views, editable action proposals, and in-view correction. It also builds procedural skills from expert videos and human teaching, then uses those skills through a Skill Agent. On LIBERO-Pro, WAA reports a 75.6% average success rate using skills evolved only from LIBERO-90. Fine-tuning Qwen3.5-9B on WAA traces raises out-of-domain success from 1.7% to 43.3%. HF Daily Papers' note

score 5

Categories: Research