One note in three: a verified census of three deployed AI scribes, and the instrument that counted it
The audit says 31.3% of the sampled AI-scribe notes contained a verified failure.
The paper tested three commercial ambient clinical scribes on 142 consultations, producing 565 notes from UK primary-care, US ambulatory, and authored scenarios. The verified failures clustered around allergy and medication information, invented patient identity, and telephone histories written as physical examinations. When identity and date errors were excluded because no product had a patient record, the failure rate was 24.8%. The authors also found that the measurement instrument changed results sharply: review instructions and model family moved verification rates enough to explain large gaps between published audits. ArXiv · AI/CL/LG's note
The paper tested three commercial ambient clinical scribes on 142 consultations, producing 565 notes from UK primary-care, US ambulatory, and authored scenarios. The verified failures clustered around allergy and medication information, invented patient identity, and telephone histories written as physical examinations. When identity and date errors were excluded because no product had a patient record, the failure rate was 24.8%. The authors also found that the measurement instrument changed results sharply: review instructions and model family moved verification rates enough to explain large gaps between published audits. ArXiv · AI/CL/LG's note
score 5