What FID Hides: Detecting, Ranking, and Diagnosing Deviations in Generative Evaluation
FID can rate unrecognizable images better than held-out real ones when only mean and covariance are matched.
Hao Chen’s paper argues that FID and KID compress generative evaluation into scalar scores that can miss distributional failures and do not show whether dispersion is shrinking or spreading. It introduces ZID, a diagnostic that separates ranking, calibrated equality testing, and signed dispersion readout. In experiments, ZID detects departures that FID leaves flat or ranks backward, including high-guidance diversity collapse in DiT-XL/2 and SiT-XL/2 sweeps. ArXiv · AI/CL/LG's note
Hao Chen’s paper argues that FID and KID compress generative evaluation into scalar scores that can miss distributional failures and do not show whether dispersion is shrinking or spreading. It introduces ZID, a diagnostic that separates ranking, calibrated equality testing, and signed dispersion readout. In experiments, ZID detects departures that FID leaves flat or ranks backward, including high-guidance diversity collapse in DiT-XL/2 and SiT-XL/2 sweeps. ArXiv · AI/CL/LG's note
score 5