Candidate supply and answer selection shape the value of LLM judging in multi-agent systems
The paper finds that LLM judges help most when the right answer is already in the pool but losing to a more common wrong one.
The authors replayed fixed candidate pools across five benchmarks to separate generation from final selection. They report that adding judge signals to answer frequency raised accuracy from 63.82% to about 70.9%. Judge reliability varied by task, generator, and how rare the correct answer was among candidates. The result is a narrower claim: better selection can preserve correct answers that multi-agent systems already generated. ArXiv · AI/CL/LG's note
The authors replayed fixed candidate pools across five benchmarks to separate generation from final selection. They report that adding judge signals to answer frequency raised accuracy from 63.82% to about 70.9%. Judge reliability varied by task, generator, and how rare the correct answer was among candidates. The result is a narrower claim: better selection can preserve correct answers that multi-agent systems already generated. ArXiv · AI/CL/LG's note
score 5