Megadose AI progress, ranked daily.

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

· HF Daily Papers ·
Only 5.4% of evaluated agent runs completed the migration and preserved behavior.

The paper introduces SWE Refactor Bench, a 20-task benchmark for whole-repository migrations across four kinds of technical debt. Its evaluation checks whether the migration actually happened, whether fixed behavioral tests still pass, and whether independent agents can find hidden regressions. Across 520 runs from 8 frontier models, 13 of the 20 tasks had no accepted solution. The best model, claude-opus-5, scored 47.0/100, with agents doing far better on build toolchain rewrites than language rewrites. Source: HF Daily Papers' note

score 6

Categories: Research