Megadose Built for builders and researchers.

NAMVIS: Next-Scale Autoregressive Multi-View Image Synthesis

· HF Daily Papers ·
NAMVIS replaces diffusion denoising with coarse-to-fine token prediction for multi-view synthesis.

The paper frames sparse-view novel view synthesis as geometry-conditioned next-scale autoregression. It predicts discrete visual tokens in a small number of scale steps, with parallel sampling across tokens and target views. Its Multi-scale Projective Pose Encoding injects camera transformations into attention at each resolution. The authors report better PSNR, SSIM, and LPIPS than evaluated diffusion baselines on Objaverse, GSO, and OmniObject3D, while running more than 3x faster under the same setting. HF Daily Papers' note

score 4

Categories: Research