World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation
W2-VLA predicts future wrist-view latents so the robot can act with contact-level context, not just current camera inputs.
The paper argues that main-view and wrist-view observations serve different roles in manipulation. W2-VLA uses current multi-view observations and a task instruction to forecast future wrist-local representations, then feeds that future-aware context into action prediction. The authors also introduce W2-CoT, a synthesis pipeline for annotations about manipulation progress, physical transition cues, and wrist-local evidence. They report gains on LIBERO, RoboTwin 2.0, and real-world tasks, with action generation above 80 Hz. HF Daily Papers' note
The paper argues that main-view and wrist-view observations serve different roles in manipulation. W2-VLA uses current multi-view observations and a task instruction to forecast future wrist-local representations, then feeds that future-aware context into action prediction. The authors also introduce W2-CoT, a synthesis pipeline for annotations about manipulation progress, physical transition cues, and wrist-local evidence. They report gains on LIBERO, RoboTwin 2.0, and real-world tasks, with action generation above 80 Hz. HF Daily Papers' note
score 5