Spurious Advantage Hidden in GRPO
GRPO can over-reward answers that were reached by guessing, not reasoning.
The paper calls this a “spurious advantage” in GRPO’s advantage estimator. It says the problem shows up in bounded-answer tasks, open-answer tasks with bounded sub-cases, and search agents with enough budget to stumble into the same answer. The authors argue this pushes policies toward guess-like behavior. They propose SIGNBALANCE, reporting similar results on open-answer math and gains on bounded-answer math and search-agent benchmarks. ArXiv · AI/CL/LG's note
The paper calls this a “spurious advantage” in GRPO’s advantage estimator. It says the problem shows up in bounded-answer tasks, open-answer tasks with bounded sub-cases, and search agents with enough budget to stumble into the same answer. The authors argue this pushes policies toward guess-like behavior. They propose SIGNBALANCE, reporting similar results on open-answer math and gains on bounded-answer math and search-agent benchmarks. ArXiv · AI/CL/LG's note
score 5