Synthetic Persona Pretraining: Alignment from Token Zero
The paper tests whether alignment works better when a model is trained into an assistant persona from the start of pretraining.
The authors add value-aligned first-person reflections to pretraining documents, then use post-training dialogue data to bind that persona to the assistant identity. In models up to 3B parameters trained on 500B tokens, they report better constitution following, stronger jailbreak robustness, and fewer misaligned choices in out-of-distribution moral dilemmas without losing capabilities. Adding the same intervention only late in pretraining was weaker, especially on value priorities and dilemma behavior. ArXiv · AI/CL/LG's note
The authors add value-aligned first-person reflections to pretraining documents, then use post-training dialogue data to bind that persona to the assistant identity. In models up to 3B parameters trained on 500B tokens, they report better constitution following, stronger jailbreak robustness, and fewer misaligned choices in out-of-distribution moral dilemmas without losing capabilities. Adding the same intervention only late in pretraining was weaker, especially on value priorities and dilemma behavior. ArXiv · AI/CL/LG's note
score 5