Megadose AI progress, ranked and analyzed.

SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents

· HF Daily Papers ·
The verified set is meant to strip out leakage and flawed tasks that made SWE-Bench Pro scores look better than they were.

The authors say SWE-Bench Pro evaluations were weakened by reward hacking and task quality problems, including leaked gold solutions, hidden evaluation information, misleading prompts, and poorly scoped tests. Their verified version adds anti-hacking safeguards and minimally fixes inconsistent benchmark instances. In their evaluations, some models performed substantially worse, which they argue means earlier SWE-Bench Pro results may have overstated real coding ability. HF Daily Papers' note

score 6

Categories: Research