PixelDense: Dense Prediction as Representation Alignment for Pixel Diffusion
PixelDense splits semantic and geometric teacher signals so pixel diffusion models can use both without gradient conflict.
The paper says dense-prediction models such as Depth Anything v2 and Metric3D v2 beat a DINOv2-only REPA baseline as alignment targets in pixel-space diffusion. A simple four-teacher sum underperforms the best geometric teacher, so PixelDense uses separate semantic and geometric projection streams plus an orthogonality penalty. The teachers are frozen during training and removed at inference. Reported gains include higher GenEval, DPG-Bench, and HPS v2.1 scores, faster baseline-peak training, and better layout/background preservation in SDEdit. HF Daily Papers' note
The paper says dense-prediction models such as Depth Anything v2 and Metric3D v2 beat a DINOv2-only REPA baseline as alignment targets in pixel-space diffusion. A simple four-teacher sum underperforms the best geometric teacher, so PixelDense uses separate semantic and geometric projection streams plus an orthogonality penalty. The teachers are frozen during training and removed at inference. Reported gains include higher GenEval, DPG-Bench, and HPS v2.1 scores, faster baseline-peak training, and better layout/background preservation in SDEdit. HF Daily Papers' note
score 4