DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation
DyPES-VLA separates shared scene dynamics from robot-specific control so one manipulation policy can run across different embodiments.
The paper says the model learns dynamics priors through future prediction on cross-embodiment data, targeting object motion, contact, and interaction-driven scene changes. Its Mixture-of-Experts action head then maps those shared representations into each robot’s native action space, avoiding manual action-format alignment. The authors report state-of-the-art results in simulation and real-world tests, including 98.0% success on LIBERO, 59.25% on RoboCasa-GR1, and 89.02% on RoboTwin 2.0. Source: HF Daily Papers' note.
The paper says the model learns dynamics priors through future prediction on cross-embodiment data, targeting object motion, contact, and interaction-driven scene changes. Its Mixture-of-Experts action head then maps those shared representations into each robot’s native action space, avoiding manual action-format alignment. The authors report state-of-the-art results in simulation and real-world tests, including 98.0% success on LIBERO, 59.25% on RoboCasa-GR1, and 89.02% on RoboTwin 2.0. Source: HF Daily Papers' note.
score 5