Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
DARA reweights sparse reward signals by how often they are active in a training batch.
The paper argues that GDPO-style reward normalization still leaves a batch-level imbalance when some objectives produce nonzero relative advantages less often. It defines “advantage energy” and links it to active-group density, then uses that relation to set inverse-square-root density weights. In experiments on tool calling and math reasoning, DARA reaches targeted compliance faster than GDPO while staying competitive on final performance. HF Daily Papers' note
The paper argues that GDPO-style reward normalization still leaves a batch-level imbalance when some objectives produce nonzero relative advantages less often. It defines “advantage energy” and links it to active-group density, then uses that relation to set inverse-square-root density weights. In experiments on tool calling and math reasoning, DARA reaches targeted compliance faster than GDPO while staying competitive on final performance. HF Daily Papers' note
score 4