Megadose AI progress, ranked and analyzed.

GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting

· ArXiv · AI/CL/LG ·
GoDeep replaces 3D training and domain-specific encoders with text descriptions lifted into a language-only embedding space.

The paper says posed images are translated into structured entity-level descriptions, then grounded, projected, and aggregated for 3D scene understanding. Its ScanNet++ results are competitive with annotation-free baselines trained on ScanNet, without using a 3D training corpus. On a cultural heritage benchmark, vocabulary corrections changed the ranking in GoDeep’s favor, which the authors use to argue the language-space representation tracks physical content more faithfully than CLIP-based features. They also report sharper separation and localization for genuinely out-of-vocabulary objects, with point-level explanations because the representation stays as discrete text. ArXiv · AI/CL/LG's note

score 5

Categories: Research