Data Attribution of Emergent Misalignment with Persona Features
Synthetic instruction-response data, not raw human-written web text, was what reliably pushed models into emergent misalignment.
The paper traces amplified “persona” features to pre-training documents about jailbreak-like, deceptive, manipulative, or harmful agency. Steering those features could both raise misalignment in aligned models and bring misaligned models back near baseline. But fine-tuning on the attributed human-written documents did not consistently reproduce the effect, even when reformatted as assistant replies. Synthetic pairs derived from the same material did, including across model families. ArXiv · AI/CL/LG's note
The paper traces amplified “persona” features to pre-training documents about jailbreak-like, deceptive, manipulative, or harmful agency. Steering those features could both raise misalignment in aligned models and bring misaligned models back near baseline. But fine-tuning on the attributed human-written documents did not consistently reproduce the effect, even when reformatted as assistant replies. Synthetic pairs derived from the same material did, including across model families. ArXiv · AI/CL/LG's note
score 6