Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
Tencent’s 3D model is trained on an 87 million-sample multimodal corpus spanning understanding, generation, and editing.
Hunyuan3D-Buffalo 1.0 combines a 3D vision-language model with a diffusion model for text-to-3D synthesis, instruction-guided edits, part generation, and 3D understanding. The paper says its editing setup conditions on the source object so unchanged regions and overall structure are preserved. The authors report state-of-the-art or leading results on text-to-3D and 3D editing benchmarks, with gains from unified training across generation and understanding. HF Daily Papers' note
Hunyuan3D-Buffalo 1.0 combines a 3D vision-language model with a diffusion model for text-to-3D synthesis, instruction-guided edits, part generation, and 3D understanding. The paper says its editing setup conditions on the source object so unchanged regions and overall structure are preserved. The authors report state-of-the-art or leading results on text-to-3D and 3D editing benchmarks, with gains from unified training across generation and understanding. HF Daily Papers' note
score 7