PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
The benchmark uses agent players, not fixed action scripts, to test whether world models can complete long interactive goals.
PlayWorld contains 171 scenarios, each tied to a specified objective. The paper evaluates models on geometry consistency, interaction fidelity, out-of-sight evolution, insight evolution, video quality, and controllability. Tests across nine current world models found them unreliable on long-horizon interactive tasks, especially spatial consistency and persistent state changes. HF Daily Papers' note
PlayWorld contains 171 scenarios, each tied to a specified objective. The paper evaluates models on geometry consistency, interaction fidelity, out-of-sight evolution, insight evolution, video quality, and controllability. Tests across nine current world models found them unreliable on long-horizon interactive tasks, especially spatial consistency and persistent state changes. HF Daily Papers' note
score 4