Megadose AI progress, ranked and analyzed.

VICBench: A Multi-Language Benchmark for Code Vulnerability Detection

· ArXiv · AI/CL/LG ·
VICBench tests whether tools can find the commits that first introduced real vulnerabilities.

The paper describes 100 verified vulnerability-inducing commits tied to 100 CVEs across 88 Python, Java, and C++ projects. The set covers 48 CWE types and includes larger fixes and inducing commits than prior benchmarks. In the authors’ evaluation, V-SZZ and LLM4SZZ reached only 33.3% to 40.1% F1, leaving substantial manual work. Source: ArXiv · AI/CL/LG's note.

score 5

Categories: Research