Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis
The paper swaps prompt selection from accuracy to empirical AUROC for imbalanced clinical diagnosis tasks.
Ranking-PE evaluates candidate prompts by whether they rank positive cases above paired negatives, rather than by per-case correctness. The authors apply that change to Pareto dominance, reflection feedback, and final prompt choice without extra model calls. On three MIMIC diseases, it improved AUROC over accuracy-based prompt evolution by 5.8 points on fine-tuned Qwen3-VL-8B and 16.2 points on MedGemma-4B. The paper also says prompt search depends on a medical-grade visual backbone and cannot substitute for one. ArXiv · AI/CL/LG's note
Ranking-PE evaluates candidate prompts by whether they rank positive cases above paired negatives, rather than by per-case correctness. The authors apply that change to Pareto dominance, reflection feedback, and final prompt choice without extra model calls. On three MIMIC diseases, it improved AUROC over accuracy-based prompt evolution by 5.8 points on fine-tuned Qwen3-VL-8B and 16.2 points on MedGemma-4B. The paper also says prompt search depends on a medical-grade visual backbone and cannot substitute for one. ArXiv · AI/CL/LG's note
score 4