World Action Modeling with Progressive Visual Planning
ProWAM uses sparse visual sub-goals instead of dense video rollouts to guide robot actions over longer tasks.
The paper says the model predicts an ordered sequence of future visual checkpoints alongside actions, giving the policy anchors during execution. Its video backbone caches those sub-goal features in one pass, so replanning only needs lighter action denoising. In evaluations, ProWAM reports state-of-the-art results on LIBERO-Plus and randomized RoboTwin, plus 70.0% zero-shot real-world success in novel scenes. Source: HF Daily Papers' note.
The paper says the model predicts an ordered sequence of future visual checkpoints alongside actions, giving the policy anchors during execution. Its video backbone caches those sub-goal features in one pass, so replanning only needs lighter action denoising. In evaluations, ProWAM reports state-of-the-art results on LIBERO-Plus and randomized RoboTwin, plus 70.0% zero-shot real-world success in novel scenes. Source: HF Daily Papers' note.
score 5