Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
RPG raises simulated robot manipulation success to 95.0% after 15 practice rounds without changing model weights.
The framework mines an offline dataset for manipulation capabilities, builds related simulation practice tasks, and uses feedback to diagnose failures. It updates reusable symbolic skills and the system prompt, then keeps only revisions that survive cross-task evaluation. At test time, a multimodal LLM uses that prompt and skill library to coordinate perception and control. After calibration and hardware adaptation, the frozen system completed all 30 physical trials across three tasks. ArXiv · AI/CL/LG's note
The framework mines an offline dataset for manipulation capabilities, builds related simulation practice tasks, and uses feedback to diagnose failures. It updates reusable symbolic skills and the system prompt, then keeps only revisions that survive cross-task evaluation. At test time, a multimodal LLM uses that prompt and skill library to coordinate perception and control. After calibration and hardware adaptation, the frozen system completed all 30 physical trials across three tasks. ArXiv · AI/CL/LG's note
score 5