Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models
Self-Geometry uses pixel correspondences as pseudo ground truth to impose multi-view constraints at test time.
The paper says current 3D vision foundation models can predict depth, pose, and pointmaps quickly, but may leave geometric inconsistencies because explicit bundle-adjustment-style constraints are costly in pretraining. Its pipeline adds multi-view and epipolar consistency losses, a view sampler based on SO(3) geodesic distance, and LoRA-based lightweight adaptation. The authors report consistent gains in pose and geometry estimation across six VFMs and four benchmarks. Source: HF Daily Papers' note.
The paper says current 3D vision foundation models can predict depth, pose, and pointmaps quickly, but may leave geometric inconsistencies because explicit bundle-adjustment-style constraints are costly in pretraining. Its pipeline adds multi-view and epipolar consistency losses, a view sampler based on SO(3) geodesic distance, and LoRA-based lightweight adaptation. The authors report consistent gains in pose and geometry estimation across six VFMs and four benchmarks. Source: HF Daily Papers' note.
score 5