Megadose AI progress, ranked and analyzed.

AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research

· HF Daily Papers ·
The paper tests whether coding agents can make real research progress when the target is open-ended, not specified in advance.

AutoWorldModel-Bench gives agents a base world model and a fixed compute budget, then measures whether they can improve it across eight game environments. The setup uses structured ground-truth entity state, keeping the focus on dynamics modeling rather than perception. Across 64 sessions, the tested frontier agents improved held-out results in all but one run, with 33 sessions showing substantial gains. In 91% of sessions, the best edit changed the model or training process rather than just tuning hyperparameters. HF Daily Papers' note

score 5

Categories: Research