Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees
The paper proposes calibrated uncertainty gates so an LLM judge can answer, retrieve evidence, or abstain while keeping accepted errors under a chosen risk level.
When the judge is not confident from its own knowledge, the system routes the case to a retrieval-augmented pass with a second threshold. The authors say the false discovery rate among accepted verdicts stays below a user-set alpha with high probability, using finite-sample Clopper-Pearson intervals. On open-domain QA benchmarks, the framework reportedly holds the target error rate while accepting more judgments than single-mode baselines. ArXiv · AI/CL/LG's note
When the judge is not confident from its own knowledge, the system routes the case to a retrieval-augmented pass with a second threshold. The authors say the false discovery rate among accepted verdicts stays below a user-set alpha with high probability, using finite-sample Clopper-Pearson intervals. On open-domain QA benchmarks, the framework reportedly holds the target error rate while accepting more judgments than single-mode baselines. ArXiv · AI/CL/LG's note
score 4