Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints
Benchmarks that pass utility checks can still be clinically unrealistic.
The paper tests revisions to a Synthea-derived care-gap benchmark used through EHR workflows and downstream operational-style processing. Its baseline data were sparse and templated, including 79.44% sampled-pair missingness and 100.0% top-three token concentration. Two deterministic revisions improved realism metrics while staying above the existing utility floor; naive densification did not fix the structural problem. ArXiv · AI/CL/LG's note
The paper tests revisions to a Synthea-derived care-gap benchmark used through EHR workflows and downstream operational-style processing. Its baseline data were sparse and templated, including 79.44% sampled-pair missingness and 100.0% top-three token concentration. Two deterministic revisions improved realism metrics while staying above the existing utility floor; naive densification did not fix the structural problem. ArXiv · AI/CL/LG's note
score 4