Megadose AI progress, ranked daily.

SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation

· HF Daily Papers ·
Token-level defensive feedback cut adaptive prompt-injection success on Qwen3.6-27B from 94.0% to 9.0%.

The paper argues that earlier defensive finetuning methods fail because they score whole outputs, leaving the model without a precise signal for which tokens are unsafe. SecOPD instead compares the model’s injected-input rollout against behavior on the clean input and uses that token-level signal for training. The authors report the strongest result against PISmith adaptive attacks, while agentic tool-calling tests showed 4.7% attack success for SecOPD versus 5.5% for Meta-SecAlign. Code and model releases are listed as available. HF Daily Papers' note

score 5

Categories: Research