Megadose AI progress, ranked and analyzed.

Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams

· ArXiv · AI/CL/LG ·
Diagram-to-code is where the benchmark says current MLLMs still break down.

Diagram-MMU tests 12 multimodal models on 3.7k scientific diagrams and 18.3k human-validated questions across six domains. The paper evaluates diagram-to-code parsing, diagram-to-code editing, and diagram question answering, including agentic variants of each task. Models handled diagram QA better than parsing or editing diagrams into code. Agentic settings usually helped parsing and editing but hurt QA, with Claude-4.6 Opus reported as improving across all three. ArXiv · AI/CL/LG's note

score 5

Categories: Research