Megadose Built for builders and researchers.

DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation

· HF Daily Papers ·
DexPolicy improves dexterous manipulation by shrinking exploration noise over training instead of leaving PPO-style action noise to drift.

The paper tests scheduled exploration in trajectory-guided RL, keeping the loss, architecture, rewards, and optimizer settings fixed. Across five YCB objects, Target success improves most for FPO and GRPO, while PPO sees a smaller gain. On a RealMan RM75 arm with an Inspire/RH56 hand, the reported real-world gains are larger, including FPO rising from 25.0% to 85.0% mean Target success. The authors argue schedules should be judged by final task success under the intended execution conditions, not by training return alone. HF Daily Papers' note

score 4

Categories: Research