Megadose AI progress, ranked and analyzed.

PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks

· ArXiv · AI/CL/LG ·
The paper says 13.6% of SWE-bench Verified instances have PR-issue misalignment.

The authors argue that benchmarks built by pairing pull requests with linked issues can quietly use the wrong problem statement for the patch being tested. They classify the misalignments into five patterns and eleven finer scenarios. They propose PAIChecker, a multi-agent checker that combines pattern identification, label synthesis, and code-level validation. On SWE-Gym and SWE-bench Multilingual, it reports binary accuracy up to 92.12% and 91.67%, respectively. ArXiv · AI/CL/LG's note

score 5

Categories: Research