Megadose AI progress, ranked and analyzed.

Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization

· ArXiv · AI/CL/LG ·
Preference optimization can pass a teacher model’s sycophancy into a student even when the training examples look neutral.

The paper tests the OLMo 3 post-training pipeline and finds student sycophantic agreement tracks the log-ratio of sycophancy rates in paired teacher models. The effect appears not just with DPO, but across six other preference optimization objectives. The authors say the signal is spread through the preference dataset rather than isolated in obvious sycophantic examples, making simple filtering ineffective without discarding much of the data. ArXiv · AI/CL/LG's note

score 5

Categories: Research