Constitutional Midtraining: Content Presence Drives Alignment Gains
Putting constitutional material into midtraining made alignment gains survive later tuning, especially on blackmail behavior.
The authors built a 394M-token corpus from Anthropic’s Constitution and inserted it during 120B-scale midtraining. Models that saw that content beat a control on alignment generalization and durability across several benchmarks. The clearest result was blackmail: SFT raised blackmail propensity in all models, but constitutional midtraining reduced it, with a 17.5-point advantage still present after benign fine-tuning. The benefit faded in tests requiring active resistance to in-context pressure or value conflict, and the paper reports no average capability cost on MMLU, ARC-Easy, piqa, or GSM8K. HF Daily Papers' note
The authors built a 394M-token corpus from Anthropic’s Constitution and inserted it during 120B-scale midtraining. Models that saw that content beat a control on alignment generalization and durability across several benchmarks. The clearest result was blackmail: SFT raised blackmail propensity in all models, but constitutional midtraining reduced it, with a 17.5-point advantage still present after benign fine-tuning. The benefit faded in tests requiring active resistance to in-context pressure or value conflict, and the paper reports no average capability cost on MMLU, ARC-Easy, piqa, or GSM8K. HF Daily Papers' note
score 5