PointWAM: 3D World Action Modeling for Dexterous Robotic Manipulation
PointWAM turns human hand-motion video into 3D point trajectories a robot can retarget.
The paper frames dexterous manipulation as forecasting how a scene and hands move together in 3D, rather than predicting from RGB frames alone. Its model takes a colored point cloud plus a language instruction, predicts scene and hand trajectories in a shared space-time frame, then maps the hand motion to robot actions. The authors report that human-video pretraining improves average DexJoCo success by 56.9 points, with scene-trajectory supervision adding 10.9 points over hands-only forecasting. They say the full system beats the prior state of the art on ten DexJoCo tasks by 11.7 points and outperforms strong VLAs on a real robot. HF Daily Papers' note
The paper frames dexterous manipulation as forecasting how a scene and hands move together in 3D, rather than predicting from RGB frames alone. Its model takes a colored point cloud plus a language instruction, predicts scene and hand trajectories in a shared space-time frame, then maps the hand motion to robot actions. The authors report that human-video pretraining improves average DexJoCo success by 56.9 points, with scene-trajectory supervision adding 10.9 points over hands-only forecasting. They say the full system beats the prior state of the art on ten DexJoCo tasks by 11.7 points and outperforms strong VLAs on a real robot. HF Daily Papers' note
score 5