Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
The model is trained to build a compact 3D scene representation from multiple views before it answers.
Imagine3D-LLM adds learnable summary tokens after image tokens and decodes them into a 3D Gaussian Splatting representation. The reconstruction loss is paired with normal next-token training, so the answer is conditioned on that internal 3D layout. The authors report stronger cross-frame correspondence in the underlying image features and better results than prior methods on spatial reasoning and 3D understanding benchmarks. ArXiv · AI/CL/LG's note
Imagine3D-LLM adds learnable summary tokens after image tokens and decodes them into a 3D Gaussian Splatting representation. The reconstruction loss is paired with normal next-token training, so the answer is conditioned on that internal 3D layout. The authors report stronger cross-frame correspondence in the underlying image features and better results than prior methods on spatial reasoning and 3D understanding benchmarks. ArXiv · AI/CL/LG's note
score 5