AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
The paper finds LLM judges hit a hard reliability ceiling on difficult tool-calling workflows.
AgentJudgeBench tests 3,808 workflow-DAG cases across six topologies and three difficulty levels. Alignment falls as tasks get harder, and drops faster when judges do not see ground truth. On hard no-ground-truth cases, all six judges cluster around 77-82% alignment regardless of scale. Rubrics help by up to 6.5 points, while chain-of-thought and temperature changes do little. HF Daily Papers' note
AgentJudgeBench tests 3,808 workflow-DAG cases across six topologies and three difficulty levels. Alignment falls as tasks get harder, and drops faster when judges do not see ground truth. On hard no-ground-truth cases, all six judges cluster around 77-82% alignment regardless of scale. Rubrics help by up to 6.5 points, while chain-of-thought and temperature changes do little. HF Daily Papers' note
score 5