Building Rome from a Single Image
The paper turns a single image into a complete 3D scene mesh, including areas the camera never saw.
The method adapts an object-focused 3D generator, Trellis 2, for full indoor and outdoor scenes. It breaks scenes into distance-aware chunks, keeps close areas detailed, and uses larger chunks for far structures like buildings. It also lifts image features into 3D while marking free space, observed surfaces, and hidden regions. The authors report stronger geometric accuracy and perceptual quality than baselines on Tanks and Temples, ScanNet++, and in-the-wild images. HF Daily Papers' note
The method adapts an object-focused 3D generator, Trellis 2, for full indoor and outdoor scenes. It breaks scenes into distance-aware chunks, keeps close areas detailed, and uses larger chunks for far structures like buildings. It also lifts image features into 3D while marking free space, observed surfaces, and hidden regions. The authors report stronger geometric accuracy and perceptual quality than baselines on Tanks and Temples, ScanNet++, and in-the-wild images. HF Daily Papers' note
score 5