Megadose AI progress, ranked and analyzed.

SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents

· ArXiv · AI/CL/LG ·
SWE-Gate finds 221 test-passing agent repairs still failed review-derived constraints.

The benchmark separates functional correctness from compliance with pull-request review constraints. It contains 303 repository-level repair instances from 75 open-source Python repositories, with distinct functional and constraint tests. In experiments across four LLM backends under one coding-agent scaffold, 644 repairs passed functional tests, but about a third missed the review requirements. The authors argue that functional-only benchmarks overstate how well agents satisfy complete repository repair specifications. ArXiv · AI/CL/LG's note

score 6

Categories: Research