PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors
PrefPI uses preference-labeled robot rollouts to push pretrained policies toward behaviors they rarely produce on their own.
The paper frames preference learning as conditional generative modeling, then uses classifier-free guidance to amplify the preference signal. Its policy-iteration loop repeats that process so small gains can accumulate into out-of-distribution behavior shifts. The authors report results across diffusion policies and the PI0.5 flow-matching VLA in simulation and real-world settings. On real hardware, object transport height rose from 10.7 cm to 19.8 cm using 150 preference-labeled trajectories. ArXiv · AI/CL/LG's note
The paper frames preference learning as conditional generative modeling, then uses classifier-free guidance to amplify the preference signal. Its policy-iteration loop repeats that process so small gains can accumulate into out-of-distribution behavior shifts. The authors report results across diffusion policies and the PI0.5 flow-matching VLA in simulation and real-world settings. On real hardware, object transport height rose from 10.7 cm to 19.8 cm using 150 preference-labeled trajectories. ArXiv · AI/CL/LG's note
score 5