Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale
Even the best evaluated system found the right vulnerable files only weakly, scoring 0.229 File F1.
The paper introduces VLoc Bench, a benchmark built from 500 real vulnerabilities across 290 repositories, six package ecosystems, and 147 CWE categories. Agents get only a CWE description plus read-only terminal access and must identify affected files in the vulnerable snapshot. On patched snapshots, they must recognize that the recorded vulnerability is gone. In the evaluation, 38.4% of tasks had no correct localization from any tested model, and some systems still reported unsupported locations after remediation. ArXiv · AI/CL/LG's note
The paper introduces VLoc Bench, a benchmark built from 500 real vulnerabilities across 290 repositories, six package ecosystems, and 147 CWE categories. Agents get only a CWE description plus read-only terminal access and must identify affected files in the vulnerable snapshot. On patched snapshots, they must recognize that the recorded vulnerability is gone. In the evaluation, 38.4% of tasks had no correct localization from any tested model, and some systems still reported unsupported locations after remediation. ArXiv · AI/CL/LG's note
score 5