GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation
The paper argues that 3D consistency improves when generation starts from geometry-native features, not appearance-only latents.
GAE reparameterizes features from a geometry foundation model into a compact latent space for generation. Its latent can be decoded into appearance, depth, cameras, and point maps, giving a conditional flow model a shared geometric state to work from. In controlled comparisons, swapping in GAE reduced FVD by 12.7% on RealEstate10K and 23.1% on DL3DV. The authors also report that camera-trajectory error was halved on RealEstate10K. HF Daily Papers' note
GAE reparameterizes features from a geometry foundation model into a compact latent space for generation. Its latent can be decoded into appearance, depth, cameras, and point maps, giving a conditional flow model a shared geometric state to work from. In controlled comparisons, swapping in GAE reduced FVD by 12.7% on RealEstate10K and 23.1% on DL3DV. The authors also report that camera-trajectory error was halved on RealEstate10K. HF Daily Papers' note
score 5