Small Language Models as Judges for Rubric-Based Reinforcement Learning
A 1.7B Qwen probe judge beat larger generative-judge baselines on rubric rewards while running far faster.
The paper tests whether small models can score responses against instance-specific rubrics for reinforcement learning. Its best setup, a Qwen3-1.7B probe judge, produced stronger criterion-level agreement than generative verdicts or yes/no logprob margins. Used as a GRPO reward model, it raised RaR-Science rubric score from 0.232 to 0.643, above an 8B generative-judge baseline at 0.594. The authors say the 8B baseline took 10.7 times more reward-judge time, and transfer tests suggest the probe kept useful criterion-level reward structure across settings. HF Daily Papers' note
The paper tests whether small models can score responses against instance-specific rubrics for reinforcement learning. Its best setup, a Qwen3-1.7B probe judge, produced stronger criterion-level agreement than generative verdicts or yes/no logprob margins. Used as a GRPO reward model, it raised RaR-Science rubric score from 0.232 to 0.643, above an 8B generative-judge baseline at 0.594. The authors say the 8B baseline took 10.7 times more reward-judge time, and transfer tests suggest the probe kept useful criterion-level reward structure across settings. HF Daily Papers' note
score 5