Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training
Benign fine-tuning weakened the backdoor, but reinforcement learning often kept what was left alive.
The paper studies a supply-chain attack on software-engineering LLM agents built from third-party models. In their Qwen2.5-Coder-7B tests, the authors say PersistBD lifted post-training attack success from 20% to 74% after SFT and to 76% after SFT plus RL, without a comparable hit to benign task performance. They identify initial backdoor strength and gradient compatibility with benign training as factors that help the behavior survive. HF Daily Papers' note
The paper studies a supply-chain attack on software-engineering LLM agents built from third-party models. In their Qwen2.5-Coder-7B tests, the authors say PersistBD lifted post-training attack success from 20% to 74% after SFT and to 76% after SFT plus RL, without a comparable hit to benign task performance. They identify initial backdoor strength and gradient compatibility with benign training as factors that help the behavior survive. HF Daily Papers' note
score 5