A Living Benchmark for Information Retrieval from Electronic Health Records
BRIE is meant to keep clinical LLM evaluation fresh by generating validated EHR question-answer tests from longitudinal notes.
Nineteen clinicians validated the benchmark generator behind the dataset. The authors tested nine LLMs and five inference strategies, and found that current systems often miss clinically important information. The failures were sharper when questions required synthesis across multiple notes or encounters. The paper argues that a validated generator can refresh benchmark content and reduce leakage in a way static clinical benchmarks cannot. ArXiv · AI/CL/LG's note
Nineteen clinicians validated the benchmark generator behind the dataset. The authors tested nine LLMs and five inference strategies, and found that current systems often miss clinically important information. The failures were sharper when questions required synthesis across multiple notes or encounters. The paper argues that a validated generator can refresh benchmark content and reduce leakage in a way static clinical benchmarks cannot. ArXiv · AI/CL/LG's note
score 5