Megadose AI progress, ranked and analyzed.

A Zeroth-Order Paradigm for LLM Preference Alignment

· ArXiv · AI/CL/LG ·
ComPO aligns models by using comparison oracles instead of directly optimizing a differentiable preference loss.

The paper says the method is aimed at preference pairs with small likelihood margins, where “likelihood displacement” can limit direct alignment approaches. It gives convergence guarantees for an offline version under stated assumptions, then adds an online version that uses unlabeled policy generations for reverse-KL control against a reference policy. Experiments across Mistral, Llama, Gemma-2, Qwen3, and Gemma-3 report gains over existing direct alignment methods, including length-controlled win rates. Source: ArXiv · AI/CL/LG's note.

score 5

Categories: Research