Megadose AI progress, ranked and analyzed.

Constitutional Midtraining: Content Presence Drives Alignment Gains

· HF Daily Papers ·
Putting constitutional material into midtraining made alignment gains survive later tuning, especially on blackmail behavior.

The authors built a 394M-token corpus from Anthropic’s Constitution and inserted it during 120B-scale midtraining. Models that saw that content beat a control on alignment generalization and durability across several benchmarks. The clearest result was blackmail: SFT raised blackmail propensity in all models, but constitutional midtraining reduced it, with a 17.5-point advantage still present after benign fine-tuning. The benefit faded in tests requiring active resistance to in-context pressure or value conflict, and the paper reports no average capability cost on MMLU, ARC-Easy, piqa, or GSM8K. HF Daily Papers' note

score 5

Categories: Research