Can Multimodal Large Language Models Understand OCT?
A new OCT benchmark finds current multimodal models still fall well short of dependable retinal image understanding.
The paper introduces OCT-Bench, built from 10,076 multiple-choice questions over 4,137 OCT images from seven public datasets. Its taxonomy tests 20 tasks across perception, cognition, and reasoning, from anatomy and lesions to treatment and prognosis. The authors evaluated 20 proprietary, open-source, and medical-domain MLLMs. They report that neither medical adaptation nor larger scale reliably improved performance across the benchmark. HF Daily Papers' note
The paper introduces OCT-Bench, built from 10,076 multiple-choice questions over 4,137 OCT images from seven public datasets. Its taxonomy tests 20 tasks across perception, cognition, and reasoning, from anatomy and lesions to treatment and prognosis. The authors evaluated 20 proprietary, open-source, and medical-domain MLLMs. They report that neither medical adaptation nor larger scale reliably improved performance across the benchmark. HF Daily Papers' note
score 4