SpeedrunBench: Challenging LLM Agents with Video Game Speedrunning
The benchmark tests whether agents can keep finding faster strategies, not just solve a game once.
SpeedrunBench evaluates frontier LLM agents across nine video games where improvement requires reflection, long-horizon planning, and exploiting learned mechanics. The authors report that agents can approach human world records in simple platformers. On longer and more complex games, they still trail human performance under practical budgets. ArXiv · AI/CL/LG's note
SpeedrunBench evaluates frontier LLM agents across nine video games where improvement requires reflection, long-horizon planning, and exploiting learned mechanics. The authors report that agents can approach human world records in simple platformers. On longer and more complex games, they still trail human performance under practical budgets. ArXiv · AI/CL/LG's note
score 5