QuoteBench: How Matched Scores Can Hide Command-Path Failures
Matched scores can mask whether an agent failed at writing a command or lost it in transit.
QuoteBench tests 56 one-shot coding-agent tasks against an execution path with an added parser boundary. Replaying the same model replies through that parser cut success by 55.4 to 73.2 percentage points. When the boundary was disclosed, six configurations recovered 30.4 to 60.7 points, while two did not recover. The paper argues command-agent evaluations should report the model setup, generation contract, execution path, operating point, and final-state validator, not just a matched score. ArXiv · AI/CL/LG's note
QuoteBench tests 56 one-shot coding-agent tasks against an execution path with an added parser boundary. Replaying the same model replies through that parser cut success by 55.4 to 73.2 percentage points. When the boundary was disclosed, six configurations recovered 30.4 to 60.7 points, while two did not recover. The paper argues command-agent evaluations should report the model setup, generation contract, execution path, operating point, and final-state validator, not just a matched score. ArXiv · AI/CL/LG's note
score 5