EmbodiedSWE: Coding Agents for Long Horizon Dexterous Robotics
The paper tests whether coding agents can turn hard robot-task solutions into training data for general robot policies.
The authors introduce EMBODIEDSWE-BENCH, a simulation benchmark covering contact-rich manipulation, deformable objects, and tasks lasting up to half an hour. Frontier coding agents can solve some long-horizon tasks and reuse prior solutions across tasks and robot embodiments, but the fixes are iterative and often instance-specific. Their EMBODIEDSWE-GEN method expands one agent solution into diverse trajectories for training a vision-language-action model. More generated demonstrations improved VLA performance, and a VLA trained only on simulated agent-generated demos completed a long-horizon task on a real robot. HF Daily Papers' note
The authors introduce EMBODIEDSWE-BENCH, a simulation benchmark covering contact-rich manipulation, deformable objects, and tasks lasting up to half an hour. Frontier coding agents can solve some long-horizon tasks and reuse prior solutions across tasks and robot embodiments, but the fixes are iterative and often instance-specific. Their EMBODIEDSWE-GEN method expands one agent solution into diverse trajectories for training a vision-language-action model. More generated demonstrations improved VLA performance, and a VLA trained only on simulated agent-generated demos completed a long-horizon task on a real robot. HF Daily Papers' note
score 5