WorldGuide: Goal-Directed Video World Model for Procedural Task Execution
WorldGuide ties action planning to video generation so each next step is chosen from the visual state it just produced.
The paper frames procedural video generation as closed-loop task execution, starting from an initial image and a goal. Its Planner predicts the next atomic action or completion, while its Executor generates the video clip for that action. The authors also introduce WorldGuide Bench, with about 59K step-annotated videos across 245 tasks. WorldGuide reports 33.33% task success on that benchmark, above MiniMax-H3’s 29.90% despite MiniMax-H3 receiving reference action plans. HF Daily Papers' note
The paper frames procedural video generation as closed-loop task execution, starting from an initial image and a goal. Its Planner predicts the next atomic action or completion, while its Executor generates the video clip for that action. The authors also introduce WorldGuide Bench, with about 59K step-annotated videos across 245 tasks. WorldGuide reports 33.33% task success on that benchmark, above MiniMax-H3’s 29.90% despite MiniMax-H3 receiving reference action plans. HF Daily Papers' note
score 4