Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks
Agentic retrieval beat standard dense retrieval, but was about 160 times slower per query.
The paper tests a ReAct-style retrieval loop that pairs LLM reasoning with retrievers for complex search tasks. It reports an 8.7-point nDCG@10 gain over standard retrieval using the same embedding model. The same pipeline was competitive on both ViDoRe v3 and BRIGHT, which the authors present as evidence of better generalization. The cost is large: 107.4 seconds per query on average, versus 0.67 seconds for standard retrieval, with 764.1K input tokens and 5.8K output tokens consumed per query. HF Daily Papers' note
The paper tests a ReAct-style retrieval loop that pairs LLM reasoning with retrievers for complex search tasks. It reports an 8.7-point nDCG@10 gain over standard retrieval using the same embedding model. The same pipeline was competitive on both ViDoRe v3 and BRIGHT, which the authors present as evidence of better generalization. The cost is large: 107.4 seconds per query on average, versus 0.67 seconds for standard retrieval, with 764.1K input tokens and 5.8K output tokens consumed per query. HF Daily Papers' note
score 5