Octrees as an Explicit 3D Language
OctLLM represents 3D shapes as sparse octree occupancy tokens while keeping the base language pathway frozen.
The paper argues that latent shape codes and coordinate text strip away useful spatial structure. Its S-Octree format shortens octree sequences by dropping selected lower-level detail while preserving position and depth cues for generation and understanding. The model adds separate trainable 3D branches for mesh tokens, with text and image tokens staying on the frozen vision-language backbone. It reports better image-to-3D FID and render-grounded captioning than ShapeLLM-Omni while matching the backbone on general language benchmarks. HF Daily Papers' note
The paper argues that latent shape codes and coordinate text strip away useful spatial structure. Its S-Octree format shortens octree sequences by dropping selected lower-level detail while preserving position and depth cues for generation and understanding. The model adds separate trainable 3D branches for mesh tokens, with text and image tokens staying on the frozen vision-language backbone. It reports better image-to-3D FID and render-grounded captioning than ShapeLLM-Omni while matching the backbone on general language benchmarks. HF Daily Papers' note
score 5