Masked Visual Actions for Unified World Modeling
A single video-model checkpoint can use masked pixel trajectories as both robot-action input and object-goal instruction.
The paper introduces Masked Visual Actions, where action is expressed as a partially revealed trajectory inside the video frame. Showing robot motion makes the model predict how the scene responds; showing desired object motion makes it infer robot behavior that could produce that result. The authors say 15 hours of masked examples from real and simulated video were enough to finetune one checkpoint across scenes and robot embodiments. In manipulation tests, its imagined rollouts were used for policy evaluation, planning, and inverse modeling. HF Daily Papers' note
The paper introduces Masked Visual Actions, where action is expressed as a partially revealed trajectory inside the video frame. Showing robot motion makes the model predict how the scene responds; showing desired object motion makes it infer robot behavior that could produce that result. The authors say 15 hours of masked examples from real and simulated video were enough to finetune one checkpoint across scenes and robot embodiments. In manipulation tests, its imagined rollouts were used for policy evaluation, planning, and inverse modeling. HF Daily Papers' note
score 5