CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?
Current coding agents topped out at 45.33% Pass@1 on building working static-analysis checkers end to end.
CheckerBench tests 300 checker-synthesis tasks drawn from 297 CVEs across 167 repositories. Each task gives agents vulnerable and fixed revisions, a pinned analysis environment, and a checker scaffold. The authors also introduce CheckerLab to rebuild submissions and measure vulnerable-fixed diagnostic contrast, patch localization, false positives, and tool use. Their results across 21 model-harness configurations point to checker development as still unreliable for current agents. HF Daily Papers' note
CheckerBench tests 300 checker-synthesis tasks drawn from 297 CVEs across 167 repositories. Each task gives agents vulnerable and fixed revisions, a pinned analysis environment, and a checker scaffold. The authors also introduce CheckerLab to rebuild submissions and measure vulnerable-fixed diagnostic contrast, patch localization, false positives, and tool use. Their results across 21 model-harness configurations point to checker development as still unreliable for current agents. HF Daily Papers' note
score 6