MIRROR: Multimodal Intelligent Radiology Reasoning and Observation Reporter
MIRROR’s main claim is less about better radiology AI than about making its outputs auditable.
The prototype separates classification, localization, and report generation so the language model can only write from labels, probabilities, and anatomical regions. The authors say that prevents the report layer from inventing findings the classifier never produced, though generated wording can still overstate details. On ChestMNIST, the classifier shows real discrimination with macro AUROC 0.729, but default thresholds miss positives for most labels. The paper argues radiology metrics should be judged against imbalance-aware baselines, because aggregate scores can reward models that effectively do nothing. ArXiv · AI/CL/LG's note
The prototype separates classification, localization, and report generation so the language model can only write from labels, probabilities, and anatomical regions. The authors say that prevents the report layer from inventing findings the classifier never produced, though generated wording can still overstate details. On ChestMNIST, the classifier shows real discrimination with macro AUROC 0.729, but default thresholds miss positives for most labels. The paper argues radiology metrics should be judged against imbalance-aware baselines, because aggregate scores can reward models that effectively do nothing. ArXiv · AI/CL/LG's note
score 4