Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits
The paper says diffusion LLM safety can be weakened by mapping and pruning sparse “safety neurons,” including across model families.
The authors report large attack-success jumps on LLaDA, Dream, and Fast-dLLM after self-pruning or transfer pruning from autoregressive predecessors. They also introduce SN-Guided Diffusion, an offline black-box jailbreak method that steers generation away from safety-triggering regions. In their tests, it transfers to open and proprietary targets, including Llama-3-8B-Instruct, Qwen2.5-7B-Instruct, and Gemini-2.5-Flash-Lite, with lower generation cost than prior frameworks. ArXiv · AI/CL/LG's note
The authors report large attack-success jumps on LLaDA, Dream, and Fast-dLLM after self-pruning or transfer pruning from autoregressive predecessors. They also introduce SN-Guided Diffusion, an offline black-box jailbreak method that steers generation away from safety-triggering regions. In their tests, it transfers to open and proprietary targets, including Llama-3-8B-Instruct, Qwen2.5-7B-Instruct, and Gemini-2.5-Flash-Lite, with lower generation cost than prior frameworks. ArXiv · AI/CL/LG's note
score 5