Douyin Multimodal Embedding Model Technical Report
Douyin says DME keeps contrastive-encoder serving efficiency while adding training-time mechanisms for finer multimodal matching.
The report describes a two-stage embedding model for large-scale search and recommendation across text, image, video, and documents. Its second stage uses latent evidence reasoning and cross-conditional reconstruction during training, without adding generation at serving time. The authors report MMEB-v2 scores of 74.8 for the 2B model and 78.4 for the 9B model, plus a 2.92% relative gain on Douyin’s offline set. They also say DME is deployed in Douyin search scenarios and produced a 0.1% Lifetime gain in an online A/B test. HF Daily Papers' note
The report describes a two-stage embedding model for large-scale search and recommendation across text, image, video, and documents. Its second stage uses latent evidence reasoning and cross-conditional reconstruction during training, without adding generation at serving time. The authors report MMEB-v2 scores of 74.8 for the 2B model and 78.4 for the 9B model, plus a 2.92% relative gain on Douyin’s offline set. They also say DME is deployed in Douyin search scenarios and produced a 0.1% Lifetime gain in an online A/B test. HF Daily Papers' note
score 5