A corrective agentic hybrid RAG and an operations-grounded evaluation for a scientific facility
APS-RAG is already deployed for staff at the Advanced Photon Source and beat a BM25 baseline on an operations QA benchmark.
The system combines dense, sparse, and knowledge-graph retrieval, then adds a corrective agent loop and ReAct tooling over MCP. The authors built APS-Bench, a 50-question dataset with auditable gold answers, to test answers against facility operations knowledge. The full corrective Agentic GraphRAG reached 70.3% strict vital-nugget recall, compared with 63.8% for naive BM25. Removing the cross-encoder reranker sharply hurt recall, while the graph channel and corrective loop helped only marginally. ArXiv · AI/CL/LG's note
The system combines dense, sparse, and knowledge-graph retrieval, then adds a corrective agent loop and ReAct tooling over MCP. The authors built APS-Bench, a 50-question dataset with auditable gold answers, to test answers against facility operations knowledge. The full corrective Agentic GraphRAG reached 70.3% strict vital-nugget recall, compared with 63.8% for naive BM25. Removing the cross-encoder reranker sharply hurt recall, while the graph channel and corrective loop helped only marginally. ArXiv · AI/CL/LG's note
score 4