Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
The paper claims a 0.9B model can add five non-text modalities while leaving text retrieval weights unchanged.
Omni-Embed-Mini maps text, speech, audio, images, video, and visually rich documents into one cosine space. Its training uses dense cascaded captions as teacher targets, embedding those captions with the frozen text backbone instead of a separate teacher model. The authors say this keeps text-side weights bit-identical and avoids regression on text retrieval, while making the model far smaller than compared open omni embedders. A 2.3B variant is reported as competitive with Gemini embedding-2 on their overall-modality average. HF Daily Papers' note
Omni-Embed-Mini maps text, speech, audio, images, video, and visually rich documents into one cosine space. Its training uses dense cascaded captions as teacher targets, embedding those captions with the frozen text backbone instead of a separate teacher model. The authors say this keeps text-side weights bit-identical and avoids regression on text retrieval, while making the model far smaller than compared open omni embedders. A 2.3B variant is reported as competitive with Gemini embedding-2 on their overall-modality average. HF Daily Papers' note
score 5