A GPU-Parallel Framework for Heterogeneous Multi-Task Reinforcement Learning
Hebero tests one policy across 40 heterogeneous robot manipulation tasks in GPU-parallel simulation.
The paper introduces an Isaac Lab benchmark for joint training and standardized evaluation at that scale. Its DGPO method uses demonstrations for dense tracking rewards and asymmetric value learning, then adds IW-ABC to ease guidance as tasks improve and weight slower-learning tasks more heavily. With 50 demonstrations per task, IW-ABC reports 90.1% mean success for state inputs, 7.8 points above the strongest listed baseline. The authors also report a visual-policy result of 93.5% mean success and four real-world Piper robot tasks transferred from simulation. HF Daily Papers' note
The paper introduces an Isaac Lab benchmark for joint training and standardized evaluation at that scale. Its DGPO method uses demonstrations for dense tracking rewards and asymmetric value learning, then adds IW-ABC to ease guidance as tasks improve and weight slower-learning tasks more heavily. With 50 demonstrations per task, IW-ABC reports 90.1% mean success for state inputs, 7.8 points above the strongest listed baseline. The authors also report a visual-policy result of 93.5% mean success and four real-world Piper robot tasks transferred from simulation. HF Daily Papers' note
score 5