Megadose AI progress, ranked and analyzed.

UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation

· HF Daily Papers ·
UI2App tests whether models can infer working app behavior from screenshots alone.

The benchmark contains 327 screenshots organized into 45 coherent multi-route web app sets. It evaluates generated artifacts for executability, navigation reachability, visual fidelity, and interaction inference. In tests on six frontier vision-language models, strong visual reconstruction did not translate into strong interaction behavior. Cross-page state was especially weak, with half the models scoring zero on that dimension. HF Daily Papers' note

score 5

Categories: Research