Megadose AI progress, ranked and analyzed.

SUP-MIMIC: A Multi-Task Clinical Diagnosis Benchmark for Evaluating LLMs' Robustness to Contradictory Evidence

· ArXiv · AI/CL/LG ·
The benchmark finds leading LLMs stumble when clinical evidence points in conflicting or non-obvious directions.

SUP-MIMIC uses MIMIC-IV-v3.1 to test basic assessment, one-to-many diagnostic divergence, and many-to-one diagnostic convergence. The authors report sharp drops on the divergence and convergence tasks compared with baseline evaluation. They argue the results show reliance on statistical shortcuts rather than causal clinical reasoning. The paper also notes a conservative tilt toward “healthy” predictions, raising missed-diagnosis concerns in medical settings. ArXiv · AI/CL/LG's note

score 4

Categories: Research