On the Diffusibility of High-Dimensional Latents
The paper says reconstruction-tuned visual encoders can make latent diffusion harder by collapsing the effective signal dimension.
The authors argue that standard velocity prediction then spends capacity fitting orthogonal noise outside the useful manifold. Their proposed fix is clean-data, or `x0`, prediction, which keeps training focused on the signal. They report consistent text-to-image gains across several strong-reconstruction encoders. Accepted to ECCV 2026. ArXiv · AI/CL/LG's note
The authors argue that standard velocity prediction then spends capacity fitting orthogonal noise outside the useful manifold. Their proposed fix is clean-data, or `x0`, prediction, which keeps training focused on the signal. They report consistent text-to-image gains across several strong-reconstruction encoders. Accepted to ECCV 2026. ArXiv · AI/CL/LG's note
score 5