Megadose AI progress, ranked and analyzed.

WorldSculpt: Generating Compositional Worlds from Grounded Videos

· HF Daily Papers ·
A single-object 3D prior is adapted to build cluttered scenes as separate object meshes in one shared world frame.

The paper targets scenes with hundreds of heavily occluded objects, where each camera view shows only partial geometry. WorldSculpt extends Pixal3D with multi-view conditioning, grounding object generation in posed video observations. The authors say it is finetuned only on single objects yet generalizes to large, occluded scenes without scene-level training. They also introduce UE-MeshyScene, a benchmark with photorealistic cluttered scenes, per-object annotations, and ground-truth meshes. HF Daily Papers' note

score 5

Categories: Research