WideSWE: Can Coding Agents Coordinate Changes Across Repositories?
WideSWE tests coding agents on 120 real cross-repository software tasks, and the best configuration solved 42.50% of them.
The benchmark is built from reviewed changes across 103 software ecosystems, split evenly between bug fixes and features. Its prompts come from related issues and pull requests, with hidden tests adapted to allow varied correct implementations while checking required behavior. Across seven agent setups, full success ranged from 10.83% to 42.50%, led by Codex CLI paired with GPT-5.6-sol. The paper says failures often came from missing required changes, leaving recognized work incomplete, or changing the right repositories without satisfying the request. HF Daily Papers' note
The benchmark is built from reviewed changes across 103 software ecosystems, split evenly between bug fixes and features. Its prompts come from related issues and pull requests, with hidden tests adapted to allow varied correct implementations while checking required behavior. Across seven agent setups, full success ranged from 10.83% to 42.50%, led by Codex CLI paired with GPT-5.6-sol. The paper says failures often came from missing required changes, leaving recognized work incomplete, or changing the right repositories without satisfying the request. HF Daily Papers' note
score 6