JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
A cheap decision-only judge came within three points of a top LLM judge on standard evaluations.
The paper compares JEV-as-a-judge with sixteen generative and reward-model judges, using blinded human adjudication. It reports similar performance on ordinary preference and evidence-grounded factuality at 0.36% of the stronger comparator’s fee. The failures were larger when judging derivations or polished but wrong answers. A cascade that accepts confident JEV verdicts and escalates uncertain ones kept 99% of the comparator’s accuracy at lower cost. HF Daily Papers' note
The paper compares JEV-as-a-judge with sixteen generative and reward-model judges, using blinded human adjudication. It reports similar performance on ordinary preference and evidence-grounded factuality at 0.36% of the stronger comparator’s fee. The failures were larger when judging derivations or polished but wrong answers. A cascade that accepts confident JEV verdicts and escalates uncertain ones kept 99% of the comparator’s accuracy at lower cost. HF Daily Papers' note
score 5