InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
InterEvolve uses execution feedback at test time to evolve reward programs for new humanoid tasks without retraining the controller.
The paper pairs an object-aware forward-backward behavioral foundation model with staged reward programs that can be revised by an LLM agent and tuned numerically. Candidate programs are verified across parallel simulations, building on a library of successful skills. The authors report that evolved programs unlock behaviors human-designed rewards missed, including long-horizon simulated tasks and autonomous runs on a physical Unitree G1 using onboard egocentric perception. HF Daily Papers' note
The paper pairs an object-aware forward-backward behavioral foundation model with staged reward programs that can be revised by an LLM agent and tuned numerically. Candidate programs are verified across parallel simulations, building on a library of successful skills. The authors report that evolved programs unlock behaviors human-designed rewards missed, including long-horizon simulated tasks and autonomous runs on a physical Unitree G1 using onboard egocentric perception. HF Daily Papers' note
score 5