QF3: Fast Flow RL with Filtered Q-Gradients
QF3 uses filtered critic gradients to train flow-based robot policies faster and keep updates near reliable replay actions.
The paper presents an online off-policy RL method for flow policies, combining flow matching with a critic action gradient passed through a one-step prediction of the flow output. Its filter applies that gradient only on action dimensions close to the replay action, where the authors say the critic and prediction are more reliable. The authors report humanoid locomotion and motion-tracking training with a 10x wall-clock speedup over FPO++, plus zero-shot transfer of scratch-trained humanoid locomotion policies to hardware. They also test QF3 as a fine-tuning method for pretrained flow-based manipulation policies on ABC-Sim and Robomimic tasks. ArXiv · AI/CL/LG's note
The paper presents an online off-policy RL method for flow policies, combining flow matching with a critic action gradient passed through a one-step prediction of the flow output. Its filter applies that gradient only on action dimensions close to the replay action, where the authors say the critic and prediction are more reliable. The authors report humanoid locomotion and motion-tracking training with a 10x wall-clock speedup over FPO++, plus zero-shot transfer of scratch-trained humanoid locomotion policies to hardware. They also test QF3 as a fine-tuning method for pretrained flow-based manipulation policies on ABC-Sim and Robomimic tasks. ArXiv · AI/CL/LG's note
score 6