Megadose Built for builders and researchers.

RAGStress: A controlled benchmark for evaluating retrieval-augmented generation under knowledge-base degradation

· ArXiv · AI/CL/LG ·
The benchmark tests RAG systems by deliberately corrupting the knowledge base and measuring what breaks.

RAGStress uses four corruption types across three severity levels, built from 57 MMLU subjects and 182,546 documents. The paper reports 52,500 model-question-condition evaluations. Its main finding is that clean retrieval can hide robustness differences, while semantic corruptions do more damage than weaker signal perturbations. The authors frame it as a controlled stress test for corrupted-KB robustness, not a general model leaderboard. ArXiv · AI/CL/LG's note

score 5

Categories: Research