Megadose AI progress, ranked and analyzed.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

· HF Daily Papers ·
Security agents look different when the meter is running.

The paper evaluates offensive Cybench tasks and defensive Splunk BOTS v1 investigations by fixed cost, not just peak success. It separates inference spend from tool spend to show where models are actually buying progress. Offensive CTF performance improves with more test-time compute, with scaled open-weight models nearing frontier proprietary systems while staying cost-competitive. Defensive SOC work depends less on raw reasoning budget and more on controlled tool use, telemetry navigation, and selective enrichment. HF Daily Papers' note

score 5

Categories: Research