HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety
Agent harness attacks succeeded even when systems recognized the risk.
HarnessRisk tests safety across six phases of agent operation, from configuration through incident recovery. Its 128 sandboxed cases pair normal user goals with adversarial instructions hidden in untrusted workflow artifacts. Across tested harnesses, models, and configurations, attack success ranged from 12.6% to 80.9%, while utility stayed high at 75.0% to 97.6%. Harness Configuration was the weakest phase, and some setups still allowed substantial attacks despite detecting risks in more than 90% of runs. HF Daily Papers' note
HarnessRisk tests safety across six phases of agent operation, from configuration through incident recovery. Its 128 sandboxed cases pair normal user goals with adversarial instructions hidden in untrusted workflow artifacts. Across tested harnesses, models, and configurations, attack success ranged from 12.6% to 80.9%, while utility stayed high at 75.0% to 97.6%. Harness Configuration was the weakest phase, and some setups still allowed substantial attacks despite detecting risks in more than 90% of runs. HF Daily Papers' note
score 5