ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision
ParEvalLayer tests whether an unfinished agent benchmark can already justify the same comparison call as the full run.
The paper frames partial scores as insufficient because early results may reflect skewed task order, missing benchmark coverage, or easy unresolved comparisons. ParEvalLayer takes paired outcomes for two agent systems and applies a prechosen comparison policy to decide whether one is better, not better, needs more evidence, or should abstain. Replaying completed public benchmarks, the authors say three benchmarks matched the full-evaluation decision after only 15% to 25% of task outcomes under the main rule, while others needed more. Reports, they argue, should include the decision rule and the number of comparisons still undecided. ArXiv · AI/CL/LG's note
The paper frames partial scores as insufficient because early results may reflect skewed task order, missing benchmark coverage, or easy unresolved comparisons. ParEvalLayer takes paired outcomes for two agent systems and applies a prechosen comparison policy to decide whether one is better, not better, needs more evidence, or should abstain. Replaying completed public benchmarks, the authors say three benchmarks matched the full-evaluation decision after only 15% to 25% of task outcomes under the main rule, while others needed more. Reports, they argue, should include the decision rule and the number of comparisons still undecided. ArXiv · AI/CL/LG's note
score 4