Verifiable by Construction: Claim-Level Evaluation of Verbatim Citation in Clinical Question Answering
Models can quote clinical sources reliably, but the quotes often do not prove the claims attached to them.
The paper tests twelve LLMs on 222 synthetic clinical questions across four clinical practice guidelines. Most models produced verbatim quotes for more than 90% of factual claims with prompting alone. The failure point was substantiation: claude-opus-5 quoted text for 98.0% of claims, but fully supported only 37.1% of them. The authors frame this as a gap for clinical QA systems that need answers clinicians can verify without opening source documents. ArXiv · AI/CL/LG's note
The paper tests twelve LLMs on 222 synthetic clinical questions across four clinical practice guidelines. Most models produced verbatim quotes for more than 90% of factual claims with prompting alone. The failure point was substantiation: claude-opus-5 quoted text for 98.0% of claims, but fully supported only 37.1% of them. The authors frame this as a gap for clinical QA systems that need answers clinicians can verify without opening source documents. ArXiv · AI/CL/LG's note
score 4