Incidental information contaminates patient notes and disrupts clinical reasoning in large language models
Frontier models folded irrelevant encounter chatter into more than a third of generated clinical notes.
The paper tested 576 patient-clinician dialogues and found small talk appeared in 35% of notes, even though overall note-quality scores barely moved. In 3.7% of frontier-model notes, the incidental remarks were misattributed or treated as clinically relevant. A separate audio test found background speech from another patient encounter leaked into nearly half of transcripts, with downstream note contamination in 5.3% of cases. The authors argue clinical systems should be tested for resistance to this kind of contamination before use. HF Daily Papers' note
The paper tested 576 patient-clinician dialogues and found small talk appeared in 35% of notes, even though overall note-quality scores barely moved. In 3.7% of frontier-model notes, the incidental remarks were misattributed or treated as clinically relevant. A separate audio test found background speech from another patient encounter leaked into nearly half of transcripts, with downstream note contamination in 5.3% of cases. The authors argue clinical systems should be tested for resistance to this kind of contamination before use. HF Daily Papers' note
score 5