Megadose AI progress, ranked and analyzed.

SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment

· ArXiv · AI/CL/LG ·
SafeEvolve trains the agent’s runtime harness and model policy together from completed safety trajectories.

The paper argues that agent safety failures can come from both the final answer and the multi-step path taken through tools or environments. SafeEvolve turns trajectory-level safety evidence into bounded, auditable updates to prompts and hierarchical skills, while also using SFT and RL to teach the policy to use those harness changes. In experiments, the authors report a better safety-utility tradeoff than baselines, including a 3x ASR reduction on AgentDojo for Qwen3.5-4B while benign utility rises from 59.79% to 61.86%. ArXiv · AI/CL/LG's note

score 5

Categories: Research