OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation
Rubric-aware distillation gives the model token-level guidance before rubric-scored RL takes over.
The paper proposes RP-OPD, where a student model learns from a teacher that can see the rubric while the student cannot. RL is then applied afterward to optimize the rubric reward directly. In tests on HealthBench, ResearchQA, and RubricHub Science, the two-stage setup scored highest among the compared post-training methods. The authors also report fewer signs of reward hacking than an SFT + RL baseline on RubricHub Science. HF Daily Papers' note
The paper proposes RP-OPD, where a student model learns from a teacher that can see the rubric while the student cannot. RL is then applied afterward to optimize the rubric reward directly. In tests on HealthBench, ResearchQA, and RubricHub Science, the two-stage setup scored highest among the compared post-training methods. The authors also report fewer signs of reward hacking than an SFT + RL baseline on RubricHub Science. HF Daily Papers' note
score 4