Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead
The paper says SWE-bench’s top ranks are now too tightly clustered to treat small score gaps as real ordering.
Its audit of 254 submissions finds the leading two Verified entries both solve 396 of 500 tasks. The top ten share most of the same wins and losses, and paired tests do not separate any adjacent top-30 Verified entries at alpha 0.05. The authors argue reports should show comparison-specific resolution and model-scaffold provenance, since scaffold choice can move scores more than the spread among leading systems. ArXiv · AI/CL/LG's note
Its audit of 254 submissions finds the leading two Verified entries both solve 396 of 500 tasks. The top ten share most of the same wins and losses, and paired tests do not separate any adjacent top-30 Verified entries at alpha 0.05. The authors argue reports should show comparison-specific resolution and model-scaffold provenance, since scaffold choice can move scores more than the spread among leading systems. ArXiv · AI/CL/LG's note
score 6