Megadose AI progress, ranked and analyzed.

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases

· ArXiv · AI/CL/LG ·
A new benchmark tests medical AI on staged, image-heavy clinical cases, and the models still fall short on fully correct diagnoses.

ClinMM-Bench includes 1,089 real-world clinical cases and 3,760 medical images across eight specialties. The authors evaluated 15 multimodal large language models on both diagnostic accuracy and reasoning quality. Proprietary models led overall, but completely correct diagnoses remained limited across the field. The paper flags five recurring failure modes, including synthesis errors, perception errors, premature closure, and visual hallucination. ArXiv · AI/CL/LG's note

score 5

Categories: Research