RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests
RealSWE tests how coding agents handle the short, casual prompts users actually send.
The paper says 88% of real SWE-chat prompts contain only a problem statement or limited context, versus 7% of SWE-bench problems. Its RealSWE benchmark builds 381 task families where the underlying fix stays the same while prompt information and style vary. Across seven LLMs, realistic inputs cut resolution rates by 6.4 percentage points on average and can reshuffle model rankings. Desired behavior and motivation helped performance; environment details and reproduction steps did not show measurable benefit. HF Daily Papers' note
The paper says 88% of real SWE-chat prompts contain only a problem statement or limited context, versus 7% of SWE-bench problems. Its RealSWE benchmark builds 381 task families where the underlying fix stays the same while prompt information and style vary. Across seven LLMs, realistic inputs cut resolution rates by 6.4 percentage points on average and can reshuffle model rankings. Desired behavior and motivation helped performance; environment details and reproduction steps did not show measurable benefit. HF Daily Papers' note
score 6