Training LLM Judges from Language Feedback via Position-Selective Self-Distillation
The paper argues that LLM judges can generalize better when self-distillation ignores feedback positions that mainly sharpen wording instead of preserving meaning.
The authors train judges from natural-language preference feedback, not just final verdict accuracy. Their method compares teacher and student entropy by token position, then masks higher entropy-shift positions during distillation. In experiments, that improves out-of-distribution results over naive self-distillation. The self-distilled judges beat outcome-supervised RL baselines by 2–9 points on evaluated subjective subcategories while staying competitive on objective ones. HF Daily Papers' note
The authors train judges from natural-language preference feedback, not just final verdict accuracy. Their method compares teacher and student entropy by token position, then masks higher entropy-shift positions during distillation. In experiments, that improves out-of-distribution results over naive self-distillation. The self-distilled judges beat outcome-supervised RL baselines by 2–9 points on evaluated subjective subcategories while staying competitive on objective ones. HF Daily Papers' note
score 4