SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence
The paper argues that figure quality needs manuscript-grounded scoring, not generic image or image-text judgment.
SciFigAlign is trained on 3,857 peer-reviewed scientific figures rated for clarity, relevance, informativeness, and structure. It uses the figure crop, caption, citing paragraphs, and light paper context, then fine-tunes CLIP and SciBERT with cross-attention and fusion. On paper-level splits, it reports a macro MAE of 0.3524 and 81.64% within-paper pairwise accuracy on a 396-item test set. The authors say ablations show citing context, manuscript grounding, and ranking supervision are critical. ArXiv · AI/CL/LG's note
SciFigAlign is trained on 3,857 peer-reviewed scientific figures rated for clarity, relevance, informativeness, and structure. It uses the figure crop, caption, citing paragraphs, and light paper context, then fine-tunes CLIP and SciBERT with cross-attention and fusion. On paper-level splits, it reports a macro MAE of 0.3524 and 81.64% within-paper pairwise accuracy on a 396-item test set. The authors say ablations show citing context, manuscript grounding, and ranking supervision are critical. ArXiv · AI/CL/LG's note
score 4