Megadose AI progress, ranked and analyzed.

DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search

· ArXiv · AI/CL/LG ·
The paper releases a fully open retrieval training recipe, plus models and data, aimed at closing a reproducibility gap.

The authors rebuilt 665M English contrastive pre-training pairs from public sources and added 1.88M supervised fine-tuning pairs with hard negatives. Their 149M-parameter DenseOn and LateOn models report 56.20 and 57.22 average nDCG@10 on BEIR, respectively. They then translate the data into eight languages, producing 2.8B pairs and training multilingual versions on mmBERT-base. The paper says LateOn generalizes better to unseen languages and scripts than the dense model, pointing to token-level matching as the reason. Source: ArXiv · AI/CL/LG's note.

score 6

Categories: Model Releases, OSS & Tools