Megadose AI progress, ranked and analyzed.

QuoteBench: How Matched Scores Can Hide Command-Path Failures

· ArXiv · AI/CL/LG ·
Matched scores can mask whether an agent failed at writing a command or lost it in transit.

QuoteBench tests 56 one-shot coding-agent tasks against an execution path with an added parser boundary. Replaying the same model replies through that parser cut success by 55.4 to 73.2 percentage points. When the boundary was disclosed, six configurations recovered 30.4 to 60.7 points, while two did not recover. The paper argues command-agent evaluations should report the model setup, generation contract, execution path, operating point, and final-state validator, not just a matched score. ArXiv · AI/CL/LG's note

score 5

Categories: Research