DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation
DexPolicy improves dexterous manipulation by shrinking exploration noise over training instead of leaving PPO-style action noise to drift.
The paper tests scheduled exploration in trajectory-guided RL, keeping the loss, architecture, rewards, and optimizer settings fixed. Across five YCB objects, Target success improves most for FPO and GRPO, while PPO sees a smaller gain. On a RealMan RM75 arm with an Inspire/RH56 hand, the reported real-world gains are larger, including FPO rising from 25.0% to 85.0% mean Target success. The authors argue schedules should be judged by final task success under the intended execution conditions, not by training return alone. HF Daily Papers' note
The paper tests scheduled exploration in trajectory-guided RL, keeping the loss, architecture, rewards, and optimizer settings fixed. Across five YCB objects, Target success improves most for FPO and GRPO, while PPO sees a smaller gain. On a RealMan RM75 arm with an Inspire/RH56 hand, the reported real-world gains are larger, including FPO rising from 25.0% to 85.0% mean Target success. The authors argue schedules should be judged by final task success under the intended execution conditions, not by training return alone. HF Daily Papers' note
score 4