Megadose Built for builders and researchers.

QuoteBench: How Matched Scores Can Hide Command-Path Failures

· HF Daily Papers ·
The paper says equal-looking agent scores can mask a transport layer breaking commands after the model writes them.

QuoteBench tests 56 one-shot Bash-agent tasks across 14 incident-derived failure families, with exact final-state validation. Replaying the same model replies through an added unescaped parser cut success by 55.4 to 73.2 percentage points. When the boundary was disclosed, six configurations recovered 30.4 to 60.7 points, while two did not recover. The authors argue evaluations should report the model setup, generation contract, execution path, operating point, and validator, not just a matched score. HF Daily Papers' note

score 5

Categories: Research