KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
Open-weight models still miss most exact Kali Linux commands without tool hints.
KaliBench tests natural-language-to-CLI translation across 8,504 query-command pairs, 1,642 Kali tools, 23 capability dimensions, and five security phases. The authors emphasize syntax-level correctness, including flags, argument order, and executable command construction. Their verification pipeline combines LLM checks, sandboxed terminal execution, and human refinement, then uses the deterministic signals for runtime-free rewards. In unrestricted evaluation, no open-weight model topped 42% exact-command accuracy, while fine-tuning and reward training lifted an 8B model to performance comparable with a 685B MoE model. ArXiv · AI/CL/LG's note
KaliBench tests natural-language-to-CLI translation across 8,504 query-command pairs, 1,642 Kali tools, 23 capability dimensions, and five security phases. The authors emphasize syntax-level correctness, including flags, argument order, and executable command construction. Their verification pipeline combines LLM checks, sandboxed terminal execution, and human refinement, then uses the deterministic signals for runtime-free rewards. In unrestricted evaluation, no open-weight model topped 42% exact-command accuracy, while fine-tuning and reward training lifted an 8B model to performance comparable with a 685B MoE model. ArXiv · AI/CL/LG's note
score 5