Megadose AI progress, ranked and analyzed.

Spurious Advantage Hidden in GRPO

· ArXiv · AI/CL/LG ·
GRPO can over-reward answers that were reached by guessing, not reasoning.

The paper calls this a “spurious advantage” in GRPO’s advantage estimator. It says the problem shows up in bounded-answer tasks, open-answer tasks with bounded sub-cases, and search agents with enough budget to stumble into the same answer. The authors argue this pushes policies toward guess-like behavior. They propose SIGNBALANCE, reporting similar results on open-answer math and gains on bounded-answer math and search-agent benchmarks. ArXiv · AI/CL/LG's note

score 5

Categories: Research