Megadose AI progress, ranked and analyzed.

PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization

· HF Daily Papers ·
The paper’s claim is that DPO can be improved by correcting bad preference signals during training, not just discarding them.

PLC-DPO routes each comparison as clean, flipped, or tied using the calibrated policy-reference margin as evidence. The authors say this helps avoid harmful updates from reversed, weak, or ambiguous preference labels. Across 57 dataset-model-benchmark cells, it reports a 60.5 mean win rate against DPO, ahead of the next-best method at 55.5. The paper is listed as Findings of EMNLP 2026, with code available. HF Daily Papers' note

score 5

Categories: Research