Megadose Built for builders and researchers.

TaoD2C-Bench: Benchmarking MLLMs for Industrial UI Code Generation Beyond Visual Fidelity

· ArXiv · AI/CL/LG ·
TaoD2C-Bench tests whether MLLMs can turn real industrial UI designs into code that satisfies implementation constraints, not just looks close.

The benchmark uses 2,861 production designs from 17 commercial platforms, with 97,652 expert annotations for components, groups, alignment, and position.
It separates required constraints from choices a model is allowed to make.
The paper evaluates eight MLLMs across code generation, requirement inference, and requirement realization, finding major gaps in meeting implementation requirements.
It also reports that strong visual reconstruction does not necessarily mean the generated code follows those requirements.
ArXiv · AI/CL/LG's note

score 5

Categories: Research