Megadose AI progress, ranked and analyzed.

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

· HF Daily Papers ·
Puffin-World tries to make 3D world generation native to one multimodal model, including physics, depth, and appearance.

The paper says the system models gravity field and latitude, depth, and image data together, instead of leaning on separate offline modules. It uses an Omni-Camera representation for varied camera motions and tasks, then propagates physical dynamics across future frames. The authors also built Puffin-16M, with 15 million vision-language-camera triplets and 1 million motion trajectories, and say they released code, models, and datasets. HF Daily Papers' note

score 5

Categories: Research