StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks
StructRL rewards robots for verified subtask progress instead of waiting for full task success.
The paper targets long-horizon vision-language-action tasks where a single command requires several dependent manipulations. StructRL breaks each task into verifiable subtasks, gives intermediate rewards only when prerequisites are complete, and scales rewards by completion pace. Tested on RoboCasa365 and LIBERO-Long with GR00T-N1.5 and pi 0.5, it outperformed the online RL baselines the authors evaluated. HF Daily Papers' note
The paper targets long-horizon vision-language-action tasks where a single command requires several dependent manipulations. StructRL breaks each task into verifiable subtasks, gives intermediate rewards only when prerequisites are complete, and scales rewards by completion pace. Tested on RoboCasa365 and LIBERO-Long with GR00T-N1.5 and pi 0.5, it outperformed the online RL baselines the authors evaluated. HF Daily Papers' note
score 5