Megadose Built for builders and researchers.

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

· ArXiv · AI/CL/LG ·
Only 5.4% of tested agent runs completed the migrations and still passed behavior checks.

The paper introduces SWE Refactor Bench, a 20-task benchmark for whole-repository migrations across four kinds of technical debt. Its evaluation checks whether the migration actually happened, whether fixed tests still pass, and whether independent agents can find hidden behavior changes. In 520 runs across eight frontier models, 13 tasks had no accepted solution, and the best model scored 47.0/100. Agents did better on build toolchain rewrites than language rewrites. Source: ArXiv · AI/CL/LG's note.

score 5

Categories: Research