Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks
Poisoning risk can swing from 3% to 80% just by choosing different poisoned examples.
The paper argues that random poison sampling understates worst-case LLM backdoor vulnerability. Its SAILS method learns to score poison sets after a few hundred finetune-and-test runs, then ranks far larger candidate pools for stronger attacks. The authors report a 30-point average gain over influence baselines, with transfer to full-scale finetuning and extensions to code, agentic, and API-only settings. HF Daily Papers' note
The paper argues that random poison sampling understates worst-case LLM backdoor vulnerability. Its SAILS method learns to score poison sets after a few hundred finetune-and-test runs, then ranks far larger candidate pools for stronger attacks. The authors report a 30-point average gain over influence baselines, with transfer to full-scale finetuning and extensions to code, agentic, and API-only settings. HF Daily Papers' note
score 5