WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report
WeChat’s team says its 2B embedding model beats a prior 8B open-source baseline on MMEB-v2.
The report introduces WeMM-Embedding, a multimodal embedding family for text, images, video, visual documents, and interleaved inputs. It comes in 2B, 4B, and 9B variants, with the 9B model reporting an 80.6 overall score on MMEB-v2. The authors say the system is already deployed across WeChat recommendation and search surfaces, including Channels, Official Accounts, Moments, and e-commerce. Model weights and code have been released, according to the abstract. HF Daily Papers' note
The report introduces WeMM-Embedding, a multimodal embedding family for text, images, video, visual documents, and interleaved inputs. It comes in 2B, 4B, and 9B variants, with the 9B model reporting an 80.6 overall score on MMEB-v2. The authors say the system is already deployed across WeChat recommendation and search surfaces, including Channels, Official Accounts, Moments, and e-commerce. Model weights and code have been released, according to the abstract. HF Daily Papers' note
score 6