JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
A cheap decision-only judge nearly matched a leading LLM judge when allowed to hand off low-confidence cases.
The paper compares JEV-as-a-judge with sixteen generative and reward-model judges, using blinded human adjudication. It reports performance within three percentage points of its strongest LLM comparator on ordinary preference and evidence-grounded factuality, at 0.36% of that comparator’s fee. The weaker cases were judgments that required checking derivations or rejecting polished but wrong answers. A frozen cascade that accepts confident JEV verdicts and escalates uncertain ones kept 99% of the comparator’s accuracy at lower cost. ArXiv · AI/CL/LG's note
The paper compares JEV-as-a-judge with sixteen generative and reward-model judges, using blinded human adjudication. It reports performance within three percentage points of its strongest LLM comparator on ordinary preference and evidence-grounded factuality, at 0.36% of that comparator’s fee. The weaker cases were judgments that required checking derivations or rejecting polished but wrong answers. A frozen cascade that accepts confident JEV verdicts and escalates uncertain ones kept 99% of the comparator’s accuracy at lower cost. ArXiv · AI/CL/LG's note
score 5