SUN: Persistent Programs For Language-Grounded Control-to-Learning-to-Real Policies
The paper says persistent, typed task semantics can carry a robot policy from language-specified control into learned real-world execution.
SUN Programs define geometric and contact relations once, then compile them into MPC costs, RL rewards, transition guards, and diagnostics. The Kuafu system synthesizes those programs from language and scene semantics, screens them with MPC, and trains stage-conditioned policies without demonstrations or manually dense rewards. Across nine tasks, it reports 82.03% macro-success, ahead of sparse-reward and Stage-BC baselines. With 500 trajectories per task, its data trains DP3 policies to 46.0% simulation success and 34.7% on physical Franka and Kinova robots. ArXiv · AI/CL/LG's note
SUN Programs define geometric and contact relations once, then compile them into MPC costs, RL rewards, transition guards, and diagnostics. The Kuafu system synthesizes those programs from language and scene semantics, screens them with MPC, and trains stage-conditioned policies without demonstrations or manually dense rewards. Across nine tasks, it reports 82.03% macro-success, ahead of sparse-reward and Stage-BC baselines. With 500 trajectories per task, its data trains DP3 policies to 46.0% simulation success and 34.7% on physical Franka and Kinova robots. ArXiv · AI/CL/LG's note
score 5