LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes and What Recovers It
The paper finds standard LLM judges largely miss missing facts in AI-written clinical notes.
Across 500 audited note pairs, judge designs scored well on added or altered content but fell near chance on omissions. On single notes, none reliably flagged omissions more often than clean notes. Detection improved when the task was reframed to first list facts from the transcript, then check whether each appears in the note. The authors report a per-fact pipeline with lower false alarms and a cheaper single-call version that detects more omissions, with limits when omitted facts are restated elsewhere. ArXiv · AI/CL/LG's note
Across 500 audited note pairs, judge designs scored well on added or altered content but fell near chance on omissions. On single notes, none reliably flagged omissions more often than clean notes. Detection improved when the task was reframed to first list facts from the transcript, then check whether each appears in the note. The authors report a per-fact pipeline with lower false alarms and a cheaper single-call version that detects more omissions, with limits when omitted facts are restated elsewhere. ArXiv · AI/CL/LG's note
score 5