GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
GameHorizon claims a 5,000-hour AAA gameplay dataset and benchmark built to test models across short and long gameplay horizons.
The suite includes an automated annotator, a dataset from 21 games played by 100 expert players, and offline plus stepwise online evaluation tracks. The benchmark is designed to test visual understanding, instruction breakdown, planning, and action control without relying only on noisy live rollouts. The authors say they evaluated 47 models with more than one million model invocations, finding clear differences in capability and task difficulty. They plan to release the dataset, annotator, and benchmark. ArXiv · AI/CL/LG's note
The suite includes an automated annotator, a dataset from 21 games played by 100 expert players, and offline plus stepwise online evaluation tracks. The benchmark is designed to test visual understanding, instruction breakdown, planning, and action control without relying only on noisy live rollouts. The authors say they evaluated 47 models with more than one million model invocations, finding clear differences in capability and task difficulty. They plan to release the dataset, annotator, and benchmark. ArXiv · AI/CL/LG's note
score 5