DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
The model predicts robot rollouts from a frame, an instruction, and explicit arm/gripper actions, with extra checks to keep motion and objects faithful.
DreamX-Phi 1.0 uses PRoPE-style geometric encoding so each arm’s commanded path remains identifiable in the generated video.
It adds a depth branch for scene geometry and uses SAM3 masks with a frozen V-JEPA teacher to preserve small object consistency during grasping.
The team says a distilled few-step student is intended for efficient deployment.
They report first place on WorldArena 2.0 Track 1 and second place on Track 2 at the time of writing.
HF Daily Papers' note
DreamX-Phi 1.0 uses PRoPE-style geometric encoding so each arm’s commanded path remains identifiable in the generated video.
It adds a depth branch for scene geometry and uses SAM3 masks with a frozen V-JEPA teacher to preserve small object consistency during grasping.
The team says a distilled few-step student is intended for efficient deployment.
They report first place on WorldArena 2.0 Track 1 and second place on Track 2 at the time of writing.
HF Daily Papers' note
score 5