Megadose AI progress, ranked and analyzed.

Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents

· ArXiv · AI/CL/LG ·
Malicious agent skills can stay in runtime context long enough to make later unsafe actions look task-relevant.

The paper proposes Defense-as-Skill, making the runtime guard itself an installable and editable skill. Its guard, SkillSonar, checks sensitive actions against the user’s task boundary and routes them to allow, replan, or confirmation decisions. The authors built SCOPE-R, a task-conditioned dataset with six risk families, 206 confirmed malicious instances, and 43 benign tasks. In repeated GLM-5 runs, SkillSonar cut in-distribution attack success from 0.482 to 0.104 and out-of-distribution success from 0.606 to 0.115. ArXiv · AI/CL/LG's note

score 5

Categories: Research