Megadose AI progress, ranked and analyzed.

RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety

· ArXiv · AI/CL/LG ·
RAG can weaken safety behavior even when the underlying model has guardrails.

The paper introduces a benchmark that separates retrieval effects from retriever quality by testing four controlled conditions. Across five open-source LLMs, the authors report that stronger benign capability can coincide with higher unsafe capability. They also find that baseline model safety does not reliably carry over once retrieval is added, including in some cases with benign retrieved documents. ArXiv · AI/CL/LG's note

score 5

Categories: Research