Sherpa: Teaching LLMs to Teach Adaptively
Sherpa trains a tutor model against simulated student outcomes, not fixed teaching rubrics.
The paper introduces a multi-turn reinforcement learning setup with LLM-based student archetypes carrying different learning preferences. Its teacher model is optimized to improve those students’ learning results, and the authors report an average 20.5-point performance gain across archetypes. On MathTutorBench, the pedagogy score rises from 52.5% to 79.2%. In human pairwise comparisons, the trained teacher is preferred over the base model 79.6% of the time. HF Daily Papers' note
The paper introduces a multi-turn reinforcement learning setup with LLM-based student archetypes carrying different learning preferences. Its teacher model is optimized to improve those students’ learning results, and the authors report an average 20.5-point performance gain across archetypes. On MathTutorBench, the pedagogy score rises from 52.5% to 79.2%. In human pairwise comparisons, the trained teacher is preferred over the base model 79.6% of the time. HF Daily Papers' note
score 4