Megadose Built for builders and researchers.

SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

· HF Daily Papers ·
The benchmark finds top coding agents still miss most first-try repairs in scientific software.

SWE-bench Science covers 119 repository-level tasks from 98 GitHub repos across 20 scientific domains. The paper groups them into issue-driven, expert-exploratory, and engineering-integration tasks. Claude Code with Opus-5 (max) is reported as the best performer, but still below 50% pass@1. The authors trace failures to weak scientific abstraction, shallow repair paths, incomplete integration, and poor generalization beyond observed cases. HF Daily Papers' note

score 6

Categories: Research