Photorealistic Novel View Synthesis of Human Faces using Next-Scale Transformers
A next-scale autoregressive model is being tested as a sharper, more view-consistent alternative for high-resolution face synthesis.
The paper adapts next-scale transformers to generate multiple novel face viewpoints in one forward pass, aiming to preserve identity, detail, and geometry. The authors train on a synthetic face dataset covering varied identities and apparel, using full-resolution task-specific images only in later training stages. They report better perceptual fidelity and cross-view coherence on human subjects, and pair the output with a transformer-based Gaussian lifting pipeline for photorealistic 3D face models. ArXiv · AI/CL/LG's note
The paper adapts next-scale transformers to generate multiple novel face viewpoints in one forward pass, aiming to preserve identity, detail, and geometry. The authors train on a synthetic face dataset covering varied identities and apparel, using full-resolution task-specific images only in later training stages. They report better perceptual fidelity and cross-view coherence on human subjects, and pair the output with a transformer-based Gaussian lifting pipeline for photorealistic 3D face models. ArXiv · AI/CL/LG's note
score 4