Megadose AI progress, ranked and analyzed.

A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

· ArXiv · AI/CL/LG ·
VITA’s India- and LMIC-focused clinical corpus beat or matched frontier models on HealthBench scoring.

The paper says VITA ranked first on 4,023 English-language HealthBench questions, scoring 51.9% of possible rubric points versus 46.1% for GPT-5.4. On a 500-question rerun with newer models and a neutral open-weight judge, VITA was statistically tied with GPT-5.5 on mean per-question score, while leading on points-weighted score and questions won. The authors attribute the result to corpus specificity: disease guidelines, India-specific antimicrobial resistance data, formulary constraints, and resource-limited care protocols. They also note a tradeoff: VITA’s accuracy and completeness held up, but its communication scores were lower. ArXiv · AI/CL/LG's note

score 5

Categories: Research