ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts
The benchmark tests whether agents can turn scientific figures into editable PowerPoint slides, not just similar-looking images.
ReFigBench uses 1,000 arXiv overview figures and evaluates reconstructions across ten model, workflow, and harness configurations. The paper says perception is still a bottleneck, and that the same model can improve or degrade depending on the harness around it. Its specialized PPTX workflow often looks better to human judges, but it also removes native connectors in every configuration. Even the strongest agent remains below the rubric ceiling, leaving fidelity versus editability as the central problem. ArXiv · AI/CL/LG's note
ReFigBench uses 1,000 arXiv overview figures and evaluates reconstructions across ten model, workflow, and harness configurations. The paper says perception is still a bottleneck, and that the same model can improve or degrade depending on the harness around it. Its specialized PPTX workflow often looks better to human judges, but it also removes native connectors in every configuration. Even the strongest agent remains below the rubric ceiling, leaving fidelity versus editability as the central problem. ArXiv · AI/CL/LG's note
score 6