PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image
The benchmark finds models can reason over selected pathology evidence but still struggle to find it inside whole-slide images.
PathAgentBench tests evidence-seeking VLMs on 1,822 TCGA whole-slide images and 17,135 diagnostic paths annotated by ten board-certified pathologists. Leading open-weight models topped 93% on multi-scale reasoning, but the best text-guided localization mean IoU stayed below 0.09. In autonomous exploration, hit rates fell sharply at higher magnification, down to 0.020. ArXiv · AI/CL/LG's note
PathAgentBench tests evidence-seeking VLMs on 1,822 TCGA whole-slide images and 17,135 diagnostic paths annotated by ten board-certified pathologists. Leading open-weight models topped 93% on multi-scale reasoning, but the best text-guided localization mean IoU stayed below 0.09. In autonomous exploration, hit rates fell sharply at higher magnification, down to 0.020. ArXiv · AI/CL/LG's note
score 5