Sherpa: Teaching LLMs to Teach Adaptively
Sherpa trains teacher models against simulated student learning outcomes, not fixed teaching rubrics.
The paper introduces a multi-turn reinforcement learning setup with LLM-based student archetypes conditioned on different learning preferences. Teacher models trained this way improved instructed students’ performance by an average of 20.5 percentage points across archetypes. On MathTutorBench, the pedagogy score rose from 52.5% to 79.2%, and human evaluators preferred the trained teacher over the base model in 79.6% of pairwise comparisons. ArXiv · AI/CL/LG's note
The paper introduces a multi-turn reinforcement learning setup with LLM-based student archetypes conditioned on different learning preferences. Teacher models trained this way improved instructed students’ performance by an average of 20.5 percentage points across archetypes. On MathTutorBench, the pedagogy score rose from 52.5% to 79.2%, and human evaluators preferred the trained teacher over the base model in 79.6% of pairwise comparisons. ArXiv · AI/CL/LG's note
score 5