Megadose AI progress, ranked and analyzed.

ClinTraceBench: Source-Verifiable Longitudinal Clinical Reasoning over EHR-Derived Dialogues

· ArXiv · AI/CL/LG ·
The benchmark finds that compact patient-history methods can drop the very longitudinal evidence clinical tasks need.

ClinTraceBench tests 385 verified MIMIC-IV-derived dialogues across nine clinical reasoning task types, with event-ID provenance for source checking. The authors compare eight history representations across four model backbones on 6,271 questions. In an injection probe, compressed memory and summary methods recovered only 0–5.3% of injected positives even when the attribution sentence was present before construction. Full context produced large gains over no context, and Haiku under full context sat on the reported cost-performance frontier. ArXiv · AI/CL/LG's note

score 5

Categories: Research