Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings
Ovis-Embedding puts text, images, video, and audio into one shared embedding space instead of stitching together separate modality encoders.
The report says the model family uses a pretrained Qwen-omni backbone, then adapts it with contrastive training and low-rank initialization. Its training mix spans text, image, video, audio, and interleaved multimodal data, with sampling meant to create stronger in-batch negatives. The authors also describe focal loss, embedding distillation, and low-rank feature decomposition for compact inference. They report state-of-the-art results across MMEB-v3, MMEB-v2, MVEB, MAEB, and RTEB. HF Daily Papers' note
The report says the model family uses a pretrained Qwen-omni backbone, then adapts it with contrastive training and low-rank initialization. Its training mix spans text, image, video, audio, and interleaved multimodal data, with sampling meant to create stronger in-batch negatives. The authors also describe focal loss, embedding distillation, and low-rank feature decomposition for compact inference. They report state-of-the-art results across MMEB-v3, MMEB-v2, MVEB, MAEB, and RTEB. HF Daily Papers' note
score 6