From Reward Signal to Visual Utility: A Controlled Audit of Medical VLM Post-Training
Accuracy gains did not reliably mean the medical VLM was using the image better.
In a controlled Qwen2.5-VL-3B audit on PMC-VQA, language-model LoRA SFT raised correct-image accuracy by 1.10 percentage points, but reduced visual-benefit events and image sensitivity. The paper reports 155 acquired and 203 lost visual-benefit events in paired records. Broader multimodal adaptation performed worse than language-model LoRA SFT on correct-image accuracy. Standard GRPO and the tested evidence objective showed optimization activity, but held-out visual gains remained inconsistent. ArXiv · AI/CL/LG's note
In a controlled Qwen2.5-VL-3B audit on PMC-VQA, language-model LoRA SFT raised correct-image accuracy by 1.10 percentage points, but reduced visual-benefit events and image sensitivity. The paper reports 155 acquired and 203 lost visual-benefit events in paired records. Broader multimodal adaptation performed worse than language-model LoRA SFT on correct-image accuracy. Standard GRPO and the tested evidence objective showed optimization activity, but held-out visual gains remained inconsistent. ArXiv · AI/CL/LG's note
score 4