EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents
EvoSafeHarness builds a safety layer tailored to a specific frozen agent model and target domain.
The paper says fixed expert-designed harnesses can either over-block useful behavior or miss domain-specific risks. EvoSafeHarness searches both a natural-language policy and executable enforcement logic, using model behavior, domain specs, and adversarial review. Across several agent benchmarks, it reports lower attack success rates while preserving more utility than baseline defenses. HF Daily Papers' note
The paper says fixed expert-designed harnesses can either over-block useful behavior or miss domain-specific risks. EvoSafeHarness searches both a natural-language policy and executable enforcement logic, using model behavior, domain specs, and adversarial review. Across several agent benchmarks, it reports lower attack success rates while preserving more utility than baseline defenses. HF Daily Papers' note
score 5