UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
UniWorld-View pairs explicit 3D guidance with video diffusion to make wide-angle novel views from sparse monocular inputs.
The paper targets cases where only limited video or image coverage is available, where NeRF and 3D Gaussian Splatting degrade and occlusions become a problem. Its method uses occlusion-aware point cloud rendering to give the diffusion model stronger geometric priors and camera control. The authors report gains on WorldScore and zero-shot novel-view benchmarks for controllability, consistency, and visual fidelity. They also say the generated multi-view videos can support downstream dynamic 3DGS reconstruction. HF Daily Papers' note
The paper targets cases where only limited video or image coverage is available, where NeRF and 3D Gaussian Splatting degrade and occlusions become a problem. Its method uses occlusion-aware point cloud rendering to give the diffusion model stronger geometric priors and camera control. The authors report gains on WorldScore and zero-shot novel-view benchmarks for controllability, consistency, and visual fidelity. They also say the generated multi-view videos can support downstream dynamic 3DGS reconstruction. HF Daily Papers' note
score 5