Megadose Built for builders and researchers.

CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning

· HF Daily Papers ·
CorrGRPO changes GRPO’s multi-reward normalization so large-scale correlated rewards do not drown out smaller signals.

The paper keeps the centered total reward but normalizes pairwise covariances as Pearson correlations. That lets advantage magnitudes still respond to reward dependence without being driven mainly by reward scale. The authors test it against GRPO and variants on code generation, tool calling, and agent security with 0.5B to 8B models. They report gains across all three domains. HF Daily Papers' note

score 4

Categories: Research