Megadose AI progress, ranked and analyzed.

SWE-Touch: Benchmarking Coding Agents When Users Touch the Code

· ArXiv · AI/CL/LG ·
Counter-edits cut coding-agent success by 7.7 points on SWE-bench Verified.

SWE-Touch tests agents in shared workspaces by injecting plausible user edits that conflict with the repair task. The authors evaluate nine coding models, then extend the setup to longer-horizon tasks from SWE-Bench Pro and DeepSWE. The paper ties the failures to weak awareness of changing repository state: agents may keep conflicting code, or overwrite it without re-inspecting and testing the affected behavior. ArXiv · AI/CL/LG's note

score 5

Categories: Research