Megadose AI progress, ranked and analyzed.

EVO-WAM: Evolving World Action Models through Video-Action Verification

· HF Daily Papers ·
The paper claims robots can improve on unseen tasks by retraining on their own generated rollouts, after filtering for both task completion and action consistency.

EVO-WAM uses a vision-language model to pick generated video prefixes that appear to complete the task, then checks whether the paired actions match the video with an inverse dynamics model. The verified prefixes are fed back into the world action model for iterative training, without testing candidate actions in an external environment. On seven unseen RoboTwin 2.0 tasks, the reported average success rate rises from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero. In three real-world long-horizon composite tasks, Cosmos3 improves from 20.0% to 76.7%. HF Daily Papers' note

score 5

Categories: Research