GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space
GeoVerse tries to keep generated camera views tied to one shared 3D scene, rather than letting each new view drift on its own.
The paper combines a pretrained 3D foundation model’s geometric latent space with appearance features drawn from Wan2.2 VACE. A ControlNet-style adapter injects those video-model priors to help fill unseen regions while preserving structure already observed. A global spatial memory aggregates observed and synthesized content, then reprojects guidance for later views. The authors report gains over GLD, including 2.23 dB higher PSNR on DL3DV and 32.4% lower ATE on Mip-NeRF360. HF Daily Papers' note
The paper combines a pretrained 3D foundation model’s geometric latent space with appearance features drawn from Wan2.2 VACE. A ControlNet-style adapter injects those video-model priors to help fill unseen regions while preserving structure already observed. A global spatial memory aggregates observed and synthesized content, then reprojects guidance for later views. The authors report gains over GLD, including 2.23 dB higher PSNR on DL3DV and 32.4% lower ATE on Mip-NeRF360. HF Daily Papers' note
score 5