A Zeroth-Order Paradigm for LLM Preference Alignment
ComPO treats preference alignment as a comparison-oracle problem instead of directly training on a differentiable preference loss.
The paper argues this can recover directional signal from preference pairs with small likelihood margins, where “likelihood displacement” can weaken direct methods. It gives convergence guarantees for an offline version under stated smoothness, sparsity, and oracle-compatibility assumptions. An online variant adds unlabeled policy generations for reverse-KL control against a reference policy. Experiments across Mistral, Llama, Gemma, Qwen3, and Gemma-3 models report gains over existing direct alignment methods, including length-controlled win rates. Source: HF Daily Papers' note.
The paper argues this can recover directional signal from preference pairs with small likelihood margins, where “likelihood displacement” can weaken direct methods. It gives convergence guarantees for an offline version under stated smoothness, sparsity, and oracle-compatibility assumptions. An online variant adds unlabeled policy generations for reverse-KL control against a reference policy. Experiments across Mistral, Llama, Gemma, Qwen3, and Gemma-3 models report gains over existing direct alignment methods, including length-controlled win rates. Source: HF Daily Papers' note.
score 5