Rubric Rewards from Item Response Theory
RRT turns rubric verdict patterns into a reward signal instead of just adding rubric points.
The paper says summed rubric scores can hide different verdict patterns behind the same reward. Its Rubric Response Theory model estimates quality from criterion difficulty and discrimination, then updates those estimates from current rollout verdicts during training. With Qwen3.5-4B, it reports a 1.7-point macro criterion gain over GRPO across four datasets, with larger gains on hard criteria. It also says adaptive criterion selection kept performance within 0.1 points of full GRPO judging while using half the criterion budget. ArXiv · AI/CL/LG's note
The paper says summed rubric scores can hide different verdict patterns behind the same reward. Its Rubric Response Theory model estimates quality from criterion difficulty and discrimination, then updates those estimates from current rollout verdicts during training. With Qwen3.5-4B, it reports a 1.7-point macro criterion gain over GRPO across four datasets, with larger gains on hard criteria. It also says adaptive criterion selection kept performance within 0.1 points of full GRPO judging while using half the criterion budget. ArXiv · AI/CL/LG's note
score 5