ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
ClinFusion claims state-of-the-art medical multimodal performance by treating clinical AI as a vision-first problem.
The paper introduces a fused vision encoder built to handle heterogeneous 2D and native 3D medical images. It also proposes MedIF-Bench and a region-of-interest-grounded evaluation method for report generation. The authors say ClinFusion beats leading open-source medical MLLMs on 20 of 24 benchmarks and surpasses GPT-5.2 and Gemini-3-Flash on 13 of 16 multimodal benchmarks. Board-certified radiologists ranked its reports highest in a blinded review, according to the paper. HF Daily Papers' note
The paper introduces a fused vision encoder built to handle heterogeneous 2D and native 3D medical images. It also proposes MedIF-Bench and a region-of-interest-grounded evaluation method for report generation. The authors say ClinFusion beats leading open-source medical MLLMs on 20 of 24 benchmarks and surpasses GPT-5.2 and Gemini-3-Flash on 13 of 16 multimodal benchmarks. Board-certified radiologists ranked its reports highest in a blinded review, according to the paper. HF Daily Papers' note
score 5