Incremental Pooled LLM Evaluation for Cost-Effective Retrieval Model Selection
The paper says pooled LLM judging can cut repeated retrieval evaluation costs while preserving system rankings.
The authors test an incremental setup where only newly retrieved documents are judged when another retrieval candidate is added. Across four benchmarks and 11 systems, the resulting rankings tracked gold-standard evaluations closely. In their financial news QA deployment, overlap between runs let them reuse 65-80% of judgments and reduce evaluation cost by up to 4.9x. ArXiv · AI/CL/LG's note
The authors test an incremental setup where only newly retrieved documents are judged when another retrieval candidate is added. Across four benchmarks and 11 systems, the resulting rankings tracked gold-standard evaluations closely. In their financial news QA deployment, overlap between runs let them reuse 65-80% of judgments and reduce evaluation cost by up to 4.9x. ArXiv · AI/CL/LG's note
score 5