Megadose AI progress, ranked and analyzed.

MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization

· ArXiv · AI/CL/LG ·
A new benchmark isolates whether visual evidence actually helps models find the right files and functions in repo issues.

MM-IssueLoc includes 652 issue-PR examples across 23 languages, annotated by image type and relevance. It compares text-only runs with runs that include images, with gold labels at both file and function level. The reported systems still miss often: the strongest agent reaches 38.96 file Acc@5 and 22.45 function Acc@10. The paper argues that strong text-heavy SWE benchmark scores do not cleanly carry over to multimodal issue localization. ArXiv · AI/CL/LG's note

score 4

Categories: Research