Megadose AI progress, ranked and analyzed.

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

· ArXiv · AI/CL/LG ·
The benchmark scores scientific figures against their manuscript context, not just their pixels.

SciFigQual-Bench covers 6,308 images from top computer-science conference papers published from 2020 to 2025. Each figure is tied to its caption, citing sentence, and surrounding manuscript context, then rated by domain experts on clarity, layout, caption fit, context relevance, and misleading risk. The authors also introduce SFQ-Agent, a staged cross-modal evaluator; its GPT-5.6-Sol setup reports the best test performance in the abstract, with 0.418 average absolute error and 93.4% consistency. ArXiv · AI/CL/LG's note

score 4

Categories: Research