Megadose AI progress, ranked and analyzed.

MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams

· ArXiv · AI/CL/LG ·
MIRA-Ev tests whether clinical models can find and relate the evidence behind an answer, not just pick the right option.

The benchmark is built from Spanish MIR licensing-exam cases re-annotated by expert clinicians. It marks span-level premises, claims, and directed support or attack relations. The paper releases parallel Spanish, English, and Basque versions, calling it the first clinical argumentation resource in Basque. Its evaluation is split into evidence sentence retrieval, argumentative component extraction, and relation classification. ArXiv · AI/CL/LG's note

score 5

Categories: Research