SUP-MIMIC: A Multi-Task Clinical Diagnosis Benchmark for Evaluating LLMs' Robustness to Contradictory Evidence
The benchmark finds leading LLMs stumble when clinical evidence points in conflicting or non-obvious directions.
SUP-MIMIC uses MIMIC-IV-v3.1 to test basic assessment, one-to-many diagnostic divergence, and many-to-one diagnostic convergence. The authors report sharp drops on the divergence and convergence tasks compared with baseline evaluation. They argue the results show reliance on statistical shortcuts rather than causal clinical reasoning. The paper also notes a conservative tilt toward “healthy” predictions, raising missed-diagnosis concerns in medical settings. ArXiv · AI/CL/LG's note
SUP-MIMIC uses MIMIC-IV-v3.1 to test basic assessment, one-to-many diagnostic divergence, and many-to-one diagnostic convergence. The authors report sharp drops on the divergence and convergence tasks compared with baseline evaluation. They argue the results show reliance on statistical shortcuts rather than causal clinical reasoning. The paper also notes a conservative tilt toward “healthy” predictions, raising missed-diagnosis concerns in medical settings. ArXiv · AI/CL/LG's note
score 4