OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
The benchmark tests whether LLMs can follow real tumor board deliberations, and current models fall well short.
OpenTumorBoard contains 611 patient cases and 19,157 discussion turns from public tumor board recordings on YouTube. It evaluates models on specialist replies inside real discussions and on full board-style simulations that reach treatment and next-step conclusions. The best tested models scored 3.43/5 against specialist answers and 2.78/5 against recorded board conclusions. The authors say finetuning and reinforcement learning improved held-out performance, and plan to release the dataset and curation pipeline. HF Daily Papers' note
OpenTumorBoard contains 611 patient cases and 19,157 discussion turns from public tumor board recordings on YouTube. It evaluates models on specialist replies inside real discussions and on full board-style simulations that reach treatment and next-step conclusions. The best tested models scored 3.43/5 against specialist answers and 2.78/5 against recorded board conclusions. The authors say finetuning and reinforcement learning improved held-out performance, and plan to release the dataset and curation pipeline. HF Daily Papers' note
score 5