Megadose AI progress, ranked and analyzed.

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

· ArXiv · AI/CL/LG ·
ROPD tries to realign compromised fine-tuned models without depending on the attacker’s prompt template.

The paper says malicious downstream data can preserve a model’s specialist skills while adding harmful behavior that appears on demand. Its proposed Routing-based On-Policy Distillation compares aligned and compromised output distributions instead of fitting to known templates. In experiments across three datasets and three base models, the authors report stronger robustness than four baseline defenses when templates mismatch, with less damage to downstream task performance. They note ROPD is still not fully immune to template shifts, but its degradation is described as much smaller than existing methods. ArXiv · AI/CL/LG's note

score 4

Categories: Research