Megadose Built for builders and researchers.

Smaller Models, Better Rejects: Preference Distillation Scaling

· HF Daily Papers ·
Smaller frozen models produced cheaper negative examples that trained stronger 7B-to-72B students than the students’ own rejects.

The paper tests preference distillation on code generation and math reasoning. It argues that useful rejects preserve task structure while being less tied to the reference policy. Mixing smaller-model and student-scale rejects improved results as the smaller-model share rose. Lower-likelihood candidate rejects beat higher-likelihood ones across every source. HF Daily Papers' note

score 4

Categories: Research