Megadose Built for builders and researchers.

Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL

· HF Daily Papers ·
DARA reweights sparse reward signals by how often they are active in a training batch.

The paper argues that GDPO-style reward normalization still leaves a batch-level imbalance when some objectives produce nonzero relative advantages less often. It defines “advantage energy” and links it to active-group density, then uses that relation to set inverse-square-root density weights. In experiments on tool calling and math reasoning, DARA reaches targeted compliance faster than GDPO while staying competitive on final performance. HF Daily Papers' note

score 4

Categories: Research