ADEPT: Accelerating Dexterity via Pre-Training and Post-Training using Reinforcement Learning
ADEPT uses pretraining plus a guarded RL transfer recipe to keep dexterous robot skills from collapsing during fine-tuning.
The paper trains a general object-reposing policy first, then uses it as a prior for longer downstream manipulation tasks. The authors say ordinary RL fine-tuning quickly damages the pretrained reposing behavior, so they add behavior-cloning distillation, critic warm-up, and conservative on-policy updates. They report sim-to-real transfer on two high-DoF robot hands, including a setup with RGB cameras and vision-based tactile sensors. ArXiv · AI/CL/LG's note
The paper trains a general object-reposing policy first, then uses it as a prior for longer downstream manipulation tasks. The authors say ordinary RL fine-tuning quickly damages the pretrained reposing behavior, so they add behavior-cloning distillation, critic warm-up, and conservative on-policy updates. They report sim-to-real transfer on two high-DoF robot hands, including a setup with RGB cameras and vision-based tactile sensors. ArXiv · AI/CL/LG's note
score 6