Megadose AI progress, ranked and analyzed.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

· ArXiv · AI/CL/LG ·
The paper proposes guided training rollouts that still optimize for the original unguided problem.

OC-GRPO gives the model privileged hints during training, such as solution prefixes, to get past cases where hard problems produce no correct samples and no reward signal. The authors call these “off-context” rollouts because the prompt used to generate them differs from the target prompt being optimized. Their fix is an importance-corrected GRPO objective meant to avoid the instability of naïvely training on guided samples. They report a 3.9% absolute average gain over vanilla GRPO on standard math reasoning benchmarks, with negligible added cost. ArXiv · AI/CL/LG's note

score 5

Categories: Research