A Balanced Data Diet: Addressing the Exploration Bottleneck in Mega-Scale RL for Robot Control
The paper’s core claim is that large robot RL runs waste scale unless training samples stay near what the policy is just able to learn.
The authors introduce Success Guided Sampling, an adaptive sampler that shifts simulated RL training toward task configurations at the edge of the policy’s current ability. They report experiments with up to more than one million parallel environments, where SGS solves multi-terrain quadruped locomotion and contact-heavy assembly tasks that prior methods do not solve. The manipulation policies are then distilled into RGB-based policies and shown transferring zero-shot to real hardware assembly tasks. Source: ArXiv · AI/CL/LG's note.
The authors introduce Success Guided Sampling, an adaptive sampler that shifts simulated RL training toward task configurations at the edge of the policy’s current ability. They report experiments with up to more than one million parallel environments, where SGS solves multi-terrain quadruped locomotion and contact-heavy assembly tasks that prior methods do not solve. The manipulation policies are then distilled into RGB-based policies and shown transferring zero-shot to real hardware assembly tasks. Source: ArXiv · AI/CL/LG's note.
score 4