Megadose AI progress, ranked and analyzed.

EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants

· HF Daily Papers ·
The benchmark shows strong turn-by-turn UI generation still breaks down over a full five-turn workflow.

EvoGenUI-Bench tests 150 five-turn interface-maintenance tasks, covering presentation, executable interaction, and tool-grounded external state. The paper evaluates generated web artifacts in a browser using screenshots, DOM/source evidence, actor traces, and runtime logs. Across eight models, the strongest reached 74.9% Turn Pass but completed only 37.3% of full episodes. Tool-grounded tasks were harder, with Adjacent Pass Retention falling to 52.4%. HF Daily Papers' note

score 5

Categories: Research