SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
The new benchmark’s best tested agent still resolves only 41.2% of its refactoring tasks.
SWE-Bench ProMax contains 170 expert-curated tasks from real commits across Python, Java, TypeScript, Go, C, C++, and Rust. The paper says each task was rewritten and reviewed to avoid flawed specs or tests, with low-complexity and narrow-scope cases removed. The remaining instances average 11.4 modified files and 261.6 lines of code, aimed at behavior-preserving cross-file refactors. HF Daily Papers' note
SWE-Bench ProMax contains 170 expert-curated tasks from real commits across Python, Java, TypeScript, Go, C, C++, and Rust. The paper says each task was rewritten and reviewed to avoid flawed specs or tests, with low-complexity and narrow-scope cases removed. The remaining instances average 11.4 modified files and 261.6 lines of code, aimed at behavior-preserving cross-file refactors. HF Daily Papers' note
score 6