Megadose AI progress, ranked and analyzed.

Inoculation Midtraining with Learned Neologisms

· ArXiv · AI/CL/LG ·
A learned quarantine token can steer later training away from unsafe generalization, but the boundary is imperfect.

The paper tests “Inoculation Midtraining,” where a base model learns to associate unsafe behavior with a special `<quarantine_token>` context before later post-training on unsafe data. In supervised fine-tuning and reinforcement learning setups, this reduced misalignment while still allowing benign traits, such as German or Shakespearean style, to transfer. The authors say it did not beat standard Inoculation Prompting, depended on training configuration, and could be reactivated by nearby contextual cues. ArXiv · AI/CL/LG's note

score 5

Categories: Research