Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI
The paper’s core test is whether a missing clinical modality causes errors that are visible or quietly unflagged.
The authors propose a reusable, model-agnostic harness for per-example and per-modality failure analysis in multimodal clinical AI. It returns a failure taxonomy, a complementarity matrix for attributing errors to modalities, and loud-versus-silent dropout rates using deployment-observable signals. In validation with planted ground truth, the framework recovered the intended modality dominance and complementary subsets. On paired MIMIC-IV echo and ECG embeddings for LVEF and an HFrEF gate, dropping echo nearly doubled error on the held-out test split. HF Daily Papers' note
The authors propose a reusable, model-agnostic harness for per-example and per-modality failure analysis in multimodal clinical AI. It returns a failure taxonomy, a complementarity matrix for attributing errors to modalities, and loud-versus-silent dropout rates using deployment-observable signals. In validation with planted ground truth, the framework recovered the intended modality dominance and complementary subsets. On paired MIMIC-IV echo and ECG embeddings for LVEF and an HFrEF gate, dropping echo nearly doubled error on the held-out test split. HF Daily Papers' note
score 4