Megadose AI progress, ranked and analyzed.

BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms

· HF Daily Papers ·
At larger corpus sizes, BM25 beat agentic and graph-based RAG in the paper’s controlled comparison.

The study varied corpus size across 28 nested tiers while keeping the questions and key relevant/adversarial documents fixed. File-System Agent led at the smallest shared tiers, but used 39 times more query tokens at the bedrock and lost ground as the search space grew. Around 10 million corpus tokens, BM25 overtook it and stayed ahead at larger shared tiers, with a margin nearing 20 points at full scale. Dense retrieval was efficient but less accurate, while graph-based RAG hit construction limits before deployment scale. HF Daily Papers' note

score 6

Categories: Research