SWE-Touch: Benchmarking Coding Agents When Users Touch the Code
Counter-edits cut coding-agent resolve rates by 7.7 points on SWE-bench Verified.
SWE-Touch tests what happens when a user changes task-relevant code while an agent is already working. The benchmark injects plausible but conflicting edits, paired with contextual user messages, at the point where agents reach the affected code. Across nine coding models, the paper finds failures tied to weak awareness of the changed workspace: agents may keep conflicting code or overwrite it without enough re-inspection and targeted testing. HF Daily Papers' note
SWE-Touch tests what happens when a user changes task-relevant code while an agent is already working. The benchmark injects plausible but conflicting edits, paired with contextual user messages, at the point where agents reach the affected code. Across nine coding models, the paper finds failures tied to weak awareness of the changed workspace: agents may keep conflicting code or overwrite it without enough re-inspection and targeted testing. HF Daily Papers' note
score 6