Distillation Defenses Easily Break After Reinforcement Learning
Reinforcement learning can undo distillation defenses that looked effective right after copying.
The paper argues that evaluating defenses only immediately after distillation misses a realistic attacker step: further RL training. In their results, simple API-obtainable data can produce reasoning gains comparable to attacks that extract full hidden traces. The authors say defenses that leak enough information to reconstruct approximate reasoning traces are likely ineffective. They point instead toward batch-level distillation defenses as a more promising direction. HF Daily Papers' note
The paper argues that evaluating defenses only immediately after distillation misses a realistic attacker step: further RL training. In their results, simple API-obtainable data can produce reasoning gains comparable to attacks that extract full hidden traces. The authors say defenses that leak enough information to reconstruct approximate reasoning traces are likely ineffective. They point instead toward batch-level distillation defenses as a more promising direction. HF Daily Papers' note
score 5