The Embedder's Dilemma: LLMs Are Better, but at What Cost?
LLMs only match top embedding models when the bill and latency are much higher.
The paper compares ten LLMs with 26 embedding models across 37 tasks and finds the best LLM, Gemini 3.1 Pro, only 0.4 points ahead of the best embedding model. LLMs do better on reasoning-heavy retrieval, while embedding models lead on classification; the two are roughly even on clustering, STS, and pair classification. The cost gap is the point: one LLM run costs up to 1,431x more than a comparable embedding model, and open LLMs are far slower on the same GPU. The authors argue for using embedding models for most embedding work and reserving LLMs for retrieval that actually needs reasoning. HF Daily Papers' note
The paper compares ten LLMs with 26 embedding models across 37 tasks and finds the best LLM, Gemini 3.1 Pro, only 0.4 points ahead of the best embedding model. LLMs do better on reasoning-heavy retrieval, while embedding models lead on classification; the two are roughly even on clustering, STS, and pair classification. The cost gap is the point: one LLM run costs up to 1,431x more than a comparable embedding model, and open LLMs are far slower on the same GPU. The authors argue for using embedding models for most embedding work and reserving LLMs for retrieval that actually needs reasoning. HF Daily Papers' note
score 5