Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation
Marigold V2 reports 16–26% AbsRel gains over the previous best on KITTI and ETH3D.
The paper revisits Marigold as a way to turn diffusion transformer image models into monocular depth estimators. It targets cheap single-step inference from pretrained flow-matching models, with quantization where needed. The authors say naive training creates artifacts, and address them with semantic feature alignment plus a two-stage fine-tuning setup using a Sinkhorn-based loss. They claim sharper depth maps on hard details like fur, foliage, and hair-thin edges, and state-of-the-art results on related dense regression tasks. HF Daily Papers' note
The paper revisits Marigold as a way to turn diffusion transformer image models into monocular depth estimators. It targets cheap single-step inference from pretrained flow-matching models, with quantization where needed. The authors say naive training creates artifacts, and address them with semantic feature alignment plus a two-stage fine-tuning setup using a Sinkhorn-based loss. They claim sharper depth maps on hard details like fur, foliage, and hair-thin edges, and state-of-the-art results on related dense regression tasks. HF Daily Papers' note
score 5