LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks
The paper tests coaching-style textual feedback as a richer training signal than scalar rewards.
Its Experiential Learning method turns judge feedback on on-policy responses into reusable “experiential knowledge.” That knowledge conditions a teacher model, then gets distilled into the policy. The authors report better results than rubric-based RL across two policy families, including stronger generalization and less reward hacking on open-ended tasks. ArXiv · AI/CL/LG's note
Its Experiential Learning method turns judge feedback on on-policy responses into reusable “experiential knowledge.” That knowledge conditions a teacher model, then gets distilled into the policy. The authors report better results than rubric-based RL across two policy families, including stronger generalization and less reward hacking on open-ended tasks. ArXiv · AI/CL/LG's note
score 5