Megadose Built for builders and researchers.

LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models

· HF Daily Papers ·
LexReward grades legal-model answers by separating style, legal elements, and reasoning chain quality.

The paper proposes rubrics for each dimension, then uses those scores to build pairwise preference data for DPO and reward-model training. Its experiments say the rubric rewards distinguish stronger and weaker legal responses, and DPO improves performance across all three dimensions. The authors also train LexRM reward models that can guide reinforcement learning without reference answers at reward time. Source: HF Daily Papers' note

score 4

Categories: Research