MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning
MetaRubric targets “vacuous credit,” where rubric judges reward criteria even when the required evidence is missing.
The paper says that failure can survive even after the relevant information is removed, and can flip a response’s GRPO advantage. MetaRubric counters it by pairing evidence-aware policy optimization with rubric updates guided by the model’s own responses. It builds counterfactual prompts by changing one task-relevant fact, then assigns credit only when the answer contains enough evidence for the criterion. The authors report gains over static-judge GRPO on PubMedQA, HealthBench-Hard, and two multimodal medical benchmarks. HF Daily Papers' note
The paper says that failure can survive even after the relevant information is removed, and can flip a response’s GRPO advantage. MetaRubric counters it by pairing evidence-aware policy optimization with rubric updates guided by the model’s own responses. It builds counterfactual prompts by changing one task-relevant fact, then assigns credit only when the answer contains enough evidence for the criterion. The authors report gains over static-judge GRPO on PubMedQA, HealthBench-Hard, and two multimodal medical benchmarks. HF Daily Papers' note
score 4