Megadose AI progress, ranked and analyzed.

3D-Aware VLMs with Implicit and Explicit Geometries

· ArXiv · AI/CL/LG ·
VLM-IE3D adds learned 3D geometry to RGB-video VLMs without requiring extra 3D input.

The paper introduces implicit geometry tokens for high-level geometric priors and explicit geometry tokens for reconstructed 3D attributes. A 3D-aware adapter fuses those representations with standard 2D visual cues. The authors report stronger results across 3D video detection, 3D visual grounding, 3D dense captioning, and spatial reasoning. The work is listed as accepted to ECCV 2026, with code and models available. ArXiv · AI/CL/LG's note

score 4

Categories: Research