Megadose AI progress, ranked daily.

V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning

· HF Daily Papers ·
The paper trains vision-language models with itemized visual, reasoning, and instruction-following rewards instead of a single answer score.

V-Rubrics breaks reference answers into atomic propositions, then scores model outputs for visual faithfulness, reasoning consistency, and instruction following. The authors build a 50,248-example training set from 17 visually grounded sources and annotate it with Gemini-3-Pro under a shared rubric protocol. Their rubric-based GRPO beats the same SFT baseline and answer-only GRPO, with the strongest gains on knowledge-heavy and visually grounded reasoning benchmarks. HF Daily Papers' note

score 5

Categories: Research