ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams
VLMs can read basic ER diagrams, but still fail on the structures database designers care about.
The paper introduces ERUnderstand, a 2,960-diagram benchmark with machine-readable labels for fine-grained ERD evaluation. State-of-the-art vision-language models recover common elements reasonably well, with F1 above 0.74. Performance falls on weak entities, multivalued attributes, and N-ary relationships, dropping as low as 0.07 F1. Reasoning-augmented models improve results by 15-25%, but remain sensitive to wording biases and diagram complexity. ArXiv · AI/CL/LG's note
The paper introduces ERUnderstand, a 2,960-diagram benchmark with machine-readable labels for fine-grained ERD evaluation. State-of-the-art vision-language models recover common elements reasonably well, with F1 above 0.74. Performance falls on weak entities, multivalued attributes, and N-ary relationships, dropping as low as 0.07 F1. Reasoning-augmented models improve results by 15-25%, but remain sensitive to wording biases and diagram complexity. ArXiv · AI/CL/LG's note
score 4