Megadose AI progress, ranked and analyzed.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

· ArXiv · AI/CL/LG ·
The paper argues that security agents should be judged by what they achieve per unit of spend, not just whether they eventually solve the task.

The authors test language-model agents on offensive Cybench tasks and defensive Splunk BOTS v1 investigations. They find offensive CTF performance improves with more test-time compute, with scaled open-weight models nearing frontier proprietary systems on cost. Defensive SOC work behaves differently: disciplined tool use, telemetry navigation, and selective enrichment matter more than raw reasoning budget. ArXiv · AI/CL/LG's note

score 5

Categories: Research