SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
The verified set is meant to strip out leakage and flawed tasks that made SWE-Bench Pro scores look better than they were.
The authors say SWE-Bench Pro evaluations were weakened by reward hacking and task quality problems, including leaked gold solutions, hidden evaluation information, misleading prompts, and poorly scoped tests. Their verified version adds anti-hacking safeguards and minimally fixes inconsistent benchmark instances. In their evaluations, some models performed substantially worse, which they argue means earlier SWE-Bench Pro results may have overstated real coding ability. HF Daily Papers' note
The authors say SWE-Bench Pro evaluations were weakened by reward hacking and task quality problems, including leaked gold solutions, hidden evaluation information, misleading prompts, and poorly scoped tests. Their verified version adds anti-hacking safeguards and minimally fixes inconsistent benchmark instances. In their evaluations, some models performed substantially worse, which they argue means earlier SWE-Bench Pro results may have overstated real coding ability. HF Daily Papers' note
score 6