Megadose AI progress, ranked and analyzed.

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

· HF Daily Papers ·
OpenART tests agent safety by changing the environment state, not the task goal.

The paper introduces a red-teaming arena with more than 10,000 validated stateful scenarios across 50 domains. Its tasks have a median of 97 tool calls, aiming at long-horizon failures that static benchmarks may miss. The proposed EMHA method evolves authorized environment states in a black-box loop and reports an 85.0% pooled attack success rate across 75 agent-model configurations. The authors also find that an agent’s runtime implementation accounts for safety variation beyond the base model. HF Daily Papers' note

score 5

Categories: Research