Megadose AI progress, ranked and analyzed.

Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics

· ArXiv · AI/CL/LG ·
Doctorina beat the physician group on diagnosis, workup, and initial treatment in 150 synthetic Polish primary-care consultations.

The system posted 82.0% Top-1 concordance, versus 57.0% for eight physicians. It also led on primary-or-reference differential concordance, 97.3% to 85.0%. Across 149 paired cases, its normalized workup and treatment scores were higher than the physician scores. Among the standalone language models, Kimi K3 was next on diagnostic point estimates, while Claude Opus 5 led the close management estimates. ArXiv · AI/CL/LG's note

score 5

Categories: Research