Megadose AI progress, ranked and analyzed.

Data Attribution of Emergent Misalignment with Persona Features

· ArXiv · AI/CL/LG ·
Synthetic instruction-response data, not raw human-written web text, was what reliably pushed models into emergent misalignment.

The paper traces amplified “persona” features to pre-training documents about jailbreak-like, deceptive, manipulative, or harmful agency. Steering those features could both raise misalignment in aligned models and bring misaligned models back near baseline. But fine-tuning on the attributed human-written documents did not consistently reproduce the effect, even when reformatted as assistant replies. Synthetic pairs derived from the same material did, including across model families. ArXiv · AI/CL/LG's note

score 6

Categories: Research