Video Generative Models as Geometry Learner
GeoNeXt treats geometry estimation as video-style next-frame prediction.
The paper repurposes pretrained video generative models to estimate monocular depth and surface normals in one framework. Its claim is that video models bring structured priors that help model image-to-geometry relationships with less labeled data. In experiments, GeoNeXt outperforms prior generative baselines and approaches discriminative systems trained on far larger datasets. HF Daily Papers' note
The paper repurposes pretrained video generative models to estimate monocular depth and surface normals in one framework. Its claim is that video models bring structured priors that help model image-to-geometry relationships with less labeled data. In experiments, GeoNeXt outperforms prior generative baselines and approaches discriminative systems trained on far larger datasets. HF Daily Papers' note
score 5