OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
OpenART tests agent safety by changing the environment state, not the task goal.
The paper introduces a red-teaming arena with more than 10,000 validated stateful scenarios across 50 domains. Its tasks have a median of 97 tool calls, aiming at long-horizon failures that static benchmarks may miss. The proposed EMHA method evolves authorized environment states in a black-box loop and reports an 85.0% pooled attack success rate across 75 agent-model configurations. The authors also find that an agent’s runtime implementation accounts for safety variation beyond the base model. HF Daily Papers' note
The paper introduces a red-teaming arena with more than 10,000 validated stateful scenarios across 50 domains. Its tasks have a median of 97 tool calls, aiming at long-horizon failures that static benchmarks may miss. The proposed EMHA method evolves authorized environment states in a black-box loop and reports an 85.0% pooled attack success rate across 75 agent-model configurations. The authors also find that an agent’s runtime implementation accounts for safety variation beyond the base model. HF Daily Papers' note
score 5