Megadose Built for builders and researchers.

Co-Evolving Robot Orchestrators and Policies through Deployment

· HF Daily Papers ·
Robo-COP updates the robot’s orchestrator and action policy together during deployment, accepting a new policy only when it verifies gains on the targeted skills.

The paper frames frozen VLA policies as the bottleneck in agentic robot systems: the VLM orchestrator can route around failures, but cannot fix them. Robo-COP uses demonstrations from its own executions to fine-tune the policy when repeated failures suggest the data can help. In ten simulated RoboLab tasks, mean held-out success rose from 64.8% to 73.8% over the frozen-policy harness; fixed-schedule fine-tuning reached 65.8%. On three real-world tasks, held-out success rose from 38.3% to 50.0%. Source: HF Daily Papers' note

score 4

Categories: Research