Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
The paper argues that security agents should be judged by what they achieve per unit of spend, not just whether they eventually solve the task.
The authors test language-model agents on offensive Cybench tasks and defensive Splunk BOTS v1 investigations. They find offensive CTF performance improves with more test-time compute, with scaled open-weight models nearing frontier proprietary systems on cost. Defensive SOC work behaves differently: disciplined tool use, telemetry navigation, and selective enrichment matter more than raw reasoning budget. ArXiv · AI/CL/LG's note
The authors test language-model agents on offensive Cybench tasks and defensive Splunk BOTS v1 investigations. They find offensive CTF performance improves with more test-time compute, with scaled open-weight models nearing frontier proprietary systems on cost. Defensive SOC work behaves differently: disciplined tool use, telemetry navigation, and selective enrichment matter more than raw reasoning budget. ArXiv · AI/CL/LG's note
score 5