UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation
UI2App tests whether models can infer working app behavior from screenshots alone.
The benchmark contains 327 screenshots organized into 45 coherent multi-route web app sets. It evaluates generated artifacts for executability, navigation reachability, visual fidelity, and interaction inference. In tests on six frontier vision-language models, strong visual reconstruction did not translate into strong interaction behavior. Cross-page state was especially weak, with half the models scoring zero on that dimension. HF Daily Papers' note
The benchmark contains 327 screenshots organized into 45 coherent multi-route web app sets. It evaluates generated artifacts for executability, navigation reachability, visual fidelity, and interaction inference. In tests on six frontier vision-language models, strong visual reconstruction did not translate into strong interaction behavior. Cross-page state was especially weak, with half the models scoring zero on that dimension. HF Daily Papers' note
score 5