Fathom-Vaidya: Advancing Medical Reasoning with Rubric-Based Rewards
A 30B medical model trained with rubric-based reinforcement learning beat GPT-5 (thinking) on HealthBench-Hard.
The paper presents a sequential training setup for medical LLMs: first diagnostic reasoning from MedBullets-derived questions, then multi-turn clinical reasoning from 5.3k synthetic scenarios. Each scenario is paired with multi-dimensional rubrics used to judge responses during reinforcement learning. The authors report more than 10% improvement on MedXpertQA and 50.1% accuracy on HealthBench-Hard. ArXiv · AI/CL/LG's note
The paper presents a sequential training setup for medical LLMs: first diagnostic reasoning from MedBullets-derived questions, then multi-turn clinical reasoning from 5.3k synthetic scenarios. Each scenario is paired with multi-dimensional rubrics used to judge responses during reinforcement learning. The authors report more than 10% improvement on MedXpertQA and 50.1% accuracy on HealthBench-Hard. ArXiv · AI/CL/LG's note
score 5