Megadose AI progress, ranked and analyzed.

Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits

· ArXiv · AI/CL/LG ·
The paper says diffusion LLM safety can be weakened by mapping and pruning sparse “safety neurons,” including across model families.

The authors report large attack-success jumps on LLaDA, Dream, and Fast-dLLM after self-pruning or transfer pruning from autoregressive predecessors. They also introduce SN-Guided Diffusion, an offline black-box jailbreak method that steers generation away from safety-triggering regions. In their tests, it transfers to open and proprietary targets, including Llama-3-8B-Instruct, Qwen2.5-7B-Instruct, and Gemini-2.5-Flash-Lite, with lower generation cost than prior frameworks. ArXiv · AI/CL/LG's note

score 5

Categories: Research