Megadose Built for builders and researchers.

GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?

· HF Daily Papers ·
The benchmark splits agent-made game work into generation, bug fixing, and multi-turn optimization.

GameXpert-Bench tests 97 creation tasks, 100 repair tasks, and 17 six-turn optimization chains. The paper says current agents can often build playable foundations and follow explicit requests. They struggle more with finding defects, checking runtime behavior, and preserving existing functionality as changes accumulate. HF Daily Papers' note

score 4

Categories: Research