World Action Learning via Interaction-Centric Spectral Latent Guidance
WING uses human egocentric video to guide robot manipulation by isolating hand-object interaction from camera motion.
The paper says the method turns interaction-centric motion into latent actions, then aligns the slow, shared temporal structure between human videos and robot behavior. It reports 99.20% average success on LIBERO, 93.80% on RoboTwin 2.0, and 57.7% on RoboCasa-GR1. The authors also say it performs strongly on four real-world manipulation tasks under varied generalization settings. HF Daily Papers' note
The paper says the method turns interaction-centric motion into latent actions, then aligns the slow, shared temporal structure between human videos and robot behavior. It reports 99.20% average success on LIBERO, 93.80% on RoboTwin 2.0, and 57.7% on RoboCasa-GR1. The authors also say it performs strongly on four real-world manipulation tasks under varied generalization settings. HF Daily Papers' note
score 5