Megadose AI progress, ranked and analyzed.

Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale

· ArXiv · AI/CL/LG ·
Even the best evaluated system found the right vulnerable files only weakly, scoring 0.229 File F1.

The paper introduces VLoc Bench, a benchmark built from 500 real vulnerabilities across 290 repositories, six package ecosystems, and 147 CWE categories. Agents get only a CWE description plus read-only terminal access and must identify affected files in the vulnerable snapshot. On patched snapshots, they must recognize that the recorded vulnerability is gone. In the evaluation, 38.4% of tasks had no correct localization from any tested model, and some systems still reported unsupported locations after remediation. ArXiv · AI/CL/LG's note

score 5

Categories: Research