SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
A new 170-task benchmark finds top coding agents still resolve less than half of large multilingual refactoring jobs.
SWE-Bench ProMax is built from real commits across Python, Java, TypeScript, Go, C, C++, and Rust. The authors say each task was curated to avoid flawed or underspecified tests, a problem they cite in prior SWE-bench work. The benchmark focuses on behavior-preserving refactors spanning multiple files, averaging 11.4 files and 261.6 lines changed per instance. In their experiments, the best model reached a 41.2% resolve rate. ArXiv · AI/CL/LG's note
SWE-Bench ProMax is built from real commits across Python, Java, TypeScript, Go, C, C++, and Rust. The authors say each task was curated to avoid flawed or underspecified tests, a problem they cite in prior SWE-bench work. The benchmark focuses on behavior-preserving refactors spanning multiple files, averaging 11.4 files and 261.6 lines changed per instance. In their experiments, the best model reached a 41.2% resolve rate. ArXiv · AI/CL/LG's note
score 6