Megadose AI progress, ranked and analyzed.

ExecCritic: Learn to Test, Test to Improve for Coding Agents

· ArXiv · AI/CL/LG ·
ExecCritic’s result turns on whether the tests are independently valid, not merely executable.

The paper separates test-writing from repair so the same agent cannot make a patch and a matching bad test. Its fail-closed harness qualifies and freezes repository-native tests before a repair agent uses their execution feedback. On SWE-bench Verified, weak generated tests dragged resolution below the no-test baseline, while stronger tests improved it. With role-specific training, the two Qwen-based agents reached 72.6%, up 11.4 points from the no-test baseline. ArXiv · AI/CL/LG's note

score 6

Categories: Research