Megadose AI progress, ranked and analyzed.

From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention

· HF Daily Papers ·
PARTS trains robot policies on the few subtasks where long-horizon manipulation breaks down.

The paper says the frozen pretrained policy keeps handling nominal actions, while selectors and verifiers trigger residual corrections and local rewards around bottlenecks. That lets the system learn from successful subtasks even when full task completions are rare. On the reported YAM and Franka tasks, complete-task success rose from 32% to 61% and from 50% to 95%, with tens of minutes of real-world RL rollouts per task on average. Humans still identify bottlenecks and reset hardware when needed, but the authors report less intervention than comparable real-world RL fine-tuning methods. HF Daily Papers' note

score 5

Categories: Research