Learning 3D Editing without Paired Supervision via Generative Prior Distillation
The paper proposes training a feed-forward 3D editing model without paired 3D edit examples.
The method distills visual, semantic, and geometric signals from foundation models into a 3D editor through differentiable rendering. It uses an image-editing model for the main view and a vision-language model on novel views to preserve identity and follow instructions. A 3D-aware distribution-matching regularizer is added to reduce geometric collapse and multi-view inconsistency. The authors report better instruction fidelity and cross-view consistency than state-of-the-art baselines. HF Daily Papers' note
The method distills visual, semantic, and geometric signals from foundation models into a 3D editor through differentiable rendering. It uses an image-editing model for the main view and a vision-language model on novel views to preserve identity and follow instructions. A 3D-aware distribution-matching regularizer is added to reduce geometric collapse and multi-view inconsistency. The authors report better instruction fidelity and cross-view consistency than state-of-the-art baselines. HF Daily Papers' note
score 5