Megadose AI progress, ranked and analyzed.

HappyWorld-Bench

· HF Daily Papers ·
HappyWorld-Bench tests whether generated worlds still hold together when agents act inside them.

The benchmark covers video, spatial, and embodied world models, with 1,138 video prompts, 300 spatial scenes, and 254 embodied test cases. It combines human A/B comparisons in HappyWorld-Arena with automated metrics for behavioral correctness. The paper evaluates 14 video models, 9 spatial systems, and 8 embodied candidates, finding reliability gaps across all three tracks. Spatial systems topped out at 70.14% placement accuracy and 73.33% edit execution, while embodied models struggled with state preservation and changed action conditions. HF Daily Papers' note

score 5

Categories: Research