Megadose AI progress, ranked daily.

A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks

· ArXiv · AI/CL/LG ·
The proposed defense learns from successful jailbreaks by storing reusable attack-pattern rules, without retraining the model.

Hu and Hooi describe a test-time framework with persistent rule memory across interactions. When a jailbreak gets through, it abstracts the wrapper method rather than the harmful subject, so the rule can apply to a broader attack family. The authors say it works through external memory and prompting, including for black-box API models. In tests across four jailbreak families and multiple models, they report lower attack success rates while preserving benign utility and avoiding rising over-refusal as memory grows. ArXiv · AI/CL/LG's note

score 4

Categories: Research