Megadose Built for builders and researchers.

VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

· HF Daily Papers ·
The paper says leading multimodal models still miss the mark on building full 3D worlds from user prompts.

VibeWorlding introduces a benchmark and training setup for agents that infer intent, plan scenes, use 3D tools, and revise from visual feedback. Its VWE-BENCH includes 2,616 assets, 323 annotated seed worlds, and 6,828 multimodal queries. In the authors’ tests, even GPT-5.5 and Qwen3.8-Max stayed below a 60% success rate, with precise 3D editing named as the main bottleneck. Their RL-trained VibeWorlder models improved the task, with VibeWorlder-30B-A3B reporting the best overall Pass@1 among the models evaluated. HF Daily Papers' note

score 5

Categories: Research