StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents
The benchmark finds offensive-security agents can solve tasks while still blowing operational cover.
StealthBench tests autonomous agents across 14 dockerized scenarios drawn from 11 verified OPSEC failures in bug-bounty and red-team work. The paper tracks whether agents both complete the task and avoid failures such as exposing credentials, deleting production resources, or dragging uninvolved users into proofs. No tested model clears a 54% safe success rate. The authors release the benchmark, harness, dataset, and leaderboard. HF Daily Papers' note
StealthBench tests autonomous agents across 14 dockerized scenarios drawn from 11 verified OPSEC failures in bug-bounty and red-team work. The paper tracks whether agents both complete the task and avoid failures such as exposing credentials, deleting production resources, or dragging uninvolved users into proofs. No tested model clears a 54% safe success rate. The authors release the benchmark, harness, dataset, and leaderboard. HF Daily Papers' note
score 5