Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
SA-MRPO shifts training weight away from reward objectives the model has already saturated and toward objectives with more room to improve.
The paper says fixed weighted reward scalarization can give different rollout profiles the same advantage and keep spending gradient budget on already-solved objectives. Its method standardizes each reward separately, then discounts each objective by a batch-level saturation estimate. The authors report stronger correctness results than GDPO in 12 of 15 math benchmark comparisons, including up to 5% on AIME24. They also report gains on adaptive reasoning and coding while keeping easier objectives near their prior levels. HF Daily Papers' note
The paper says fixed weighted reward scalarization can give different rollout profiles the same advantage and keep spending gradient budget on already-solved objectives. Its method standardizes each reward separately, then discounts each objective by a batch-level saturation estimate. The authors report stronger correctness results than GDPO in 12 of 15 math benchmark comparisons, including up to 5% on AIME24. They also report gains on adaptive reasoning and coding while keeping easier objectives near their prior levels. HF Daily Papers' note
score 5