ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
ClinFusion claims stronger medical multimodal results by centering the model and evaluation around 2D and 3D clinical images.
The paper introduces a fused vision encoder meant to handle heterogeneous medical imaging, including native 3D inputs. It also proposes MedIF-Bench and a region-of-interest-grounded report metric aimed at matching radiologists’ clinical judgments. The authors report state-of-the-art results on most of the tested medical multimodal benchmarks, including comparisons with open-source systems and proprietary models. A blinded review by board-certified radiologists ranked ClinFusion’s reports highest. ArXiv · AI/CL/LG's note
The paper introduces a fused vision encoder meant to handle heterogeneous medical imaging, including native 3D inputs. It also proposes MedIF-Bench and a region-of-interest-grounded report metric aimed at matching radiologists’ clinical judgments. The authors report state-of-the-art results on most of the tested medical multimodal benchmarks, including comparisons with open-source systems and proprietary models. A blinded review by board-certified radiologists ranked ClinFusion’s reports highest. ArXiv · AI/CL/LG's note
score 5