The RAT: A Unified Bayesian Model for RAG Evaluation
The paper argues that RAG evaluation should model retrieval, abstention, and answer correctness together, rather than score the pipeline only at the end.
Its Bayesian framework separates whether the user got a correct answer from whether the generator acted appropriately given what retrieval found. The authors test it on 27 RAG configurations spanning three datasets, retrievers, and generators. They report that systems looking similar under marginal metrics show different behavior under the conditional breakdown. The model also treats LLM-as-judge labels as calibrated noisy observations alongside limited human judgments. ArXiv · AI/CL/LG's note
Its Bayesian framework separates whether the user got a correct answer from whether the generator acted appropriately given what retrieval found. The authors test it on 27 RAG configurations spanning three datasets, retrievers, and generators. They report that systems looking similar under marginal metrics show different behavior under the conditional breakdown. The model also treats LLM-as-judge labels as calibrated noisy observations alongside limited human judgments. ArXiv · AI/CL/LG's note
score 5