When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIP
Closing CLIP’s image-text gap can make zero-shot predictions collapse toward a few classes.
The paper argues that average cross-modal alignment is not enough to predict zero-shot accuracy. In its analysis, gap correction can reshape class decision margins so that outputs concentrate on a small subset of labels, a failure mode the authors call prediction-level hubness. Experiments across multiple datasets link accuracy drops after gap correction to that increased prediction concentration, including both linear and learned correction methods. ArXiv · AI/CL/LG's note
The paper argues that average cross-modal alignment is not enough to predict zero-shot accuracy. In its analysis, gap correction can reshape class decision margins so that outputs concentrate on a small subset of labels, a failure mode the authors call prediction-level hubness. Experiments across multiple datasets link accuracy drops after gap correction to that increased prediction concentration, including both linear and learned correction methods. ArXiv · AI/CL/LG's note
score 4