Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis
SNAP argues the weak link in encoder-based novel-view synthesis is an overpowered decoder, not the 3D signal.
The paper says spatially expressive decoders can let the model solve reconstruction while leaving the scene encoder with less transferable geometry. Its SNAP model uses a pose-conditioned local decoder and reconstructs in latent space instead of chasing low-level pixels. The authors report competitive results across localization, pose estimation, correspondence, depth, and robot manipulation, with patch features showing viewpoint invariance close to heavily supervised models. ArXiv · AI/CL/LG's note
The paper says spatially expressive decoders can let the model solve reconstruction while leaving the scene encoder with less transferable geometry. Its SNAP model uses a pose-conditioned local decoder and reconstructs in latent space instead of chasing low-level pixels. The authors report competitive results across localization, pose estimation, correspondence, depth, and robot manipulation, with patch features showing viewpoint invariance close to heavily supervised models. ArXiv · AI/CL/LG's note
score 4