Megadose AI progress, ranked and analyzed.

PoEM: Predicting RL Outcomes from Existing Policies

· ArXiv · AI/CL/LG ·
The paper claims RL results for a new reward can often be approximated from already post-trained policies, without another RL run.

PoEM combines existing post-trained models in log-policy space to predict the policy a new reward would produce. The authors show an exact relationship when the new reward is a linear combination of prior rewards, then report that approximate low-rank structure often appears even beyond that case. Its coefficients can be estimated from reward or policy outputs on samples. They validate the method on synthetic and real rewards across text and image settings. ArXiv · AI/CL/LG's note

score 5

Categories: Research