Megadose Built for builders and researchers.

EvoRiskBench: An Evolving Benchmark for Runtime Security Risks in Workspace Agents

· ArXiv · AI/CL/LG ·
A new benchmark reports attack success rates as high as 68.44% against tested workspace-agent setups.

EvoRiskBench defines runtime security cases by pairing risk entry points with technical effects through agent-mediated paths. The authors built 450 adversarial tasks across six scenarios and ran them in isolated environments with outcome checks from traces and environment state. They tested nine model-harness combinations across GPT-5.6 Sol, DeepSeek-V4-Pro-0813, Claude Opus 5, Claude Code, Codex, and OpenClaw. The paper says model choice drove more ASR variation than harness choice, and that the benchmark will be released after safety and reproducibility checks. ArXiv · AI/CL/LG's note

score 6

Categories: Research