Does Native 3D Texture Generation Necessarily Require 3D Assets for Training?
Tex-Zero is trained only on constructed image data, yet is reported to texture real 3D assets with high fidelity.
The paper argues that fine-grained color information matters more than real 3D geometry for training native 3D texture generation. Its method turns 2D images into 3D training samples by placing them on planes, then using patch-wise random rotations and aggregation to build more complex structures. The authors train both a VAE and a DiT this way, without real 3D assets in the training data. Experiments in the abstract are described as showing detailed 3D texture generation from image-only training data. HF Daily Papers' note
The paper argues that fine-grained color information matters more than real 3D geometry for training native 3D texture generation. Its method turns 2D images into 3D training samples by placing them on planes, then using patch-wise random rotations and aggregation to build more complex structures. The authors train both a VAE and a DiT this way, without real 3D assets in the training data. Experiments in the abstract are described as showing detailed 3D texture generation from image-only training data. HF Daily Papers' note
score 4