SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
The benchmark finds top coding agents still miss most first-try repairs in scientific software.
SWE-bench Science covers 119 repository-level tasks from 98 GitHub repos across 20 scientific domains. The paper groups them into issue-driven, expert-exploratory, and engineering-integration tasks. Claude Code with Opus-5 (max) is reported as the best performer, but still below 50% pass@1. The authors trace failures to weak scientific abstraction, shallow repair paths, incomplete integration, and poor generalization beyond observed cases. HF Daily Papers' note
SWE-bench Science covers 119 repository-level tasks from 98 GitHub repos across 20 scientific domains. The paper groups them into issue-driven, expert-exploratory, and engineering-integration tasks. Claude Code with Opus-5 (max) is reported as the best performer, but still below 50% pass@1. The authors trace failures to weak scientific abstraction, shallow repair paths, incomplete integration, and poor generalization beyond observed cases. HF Daily Papers' note
score 6