Megadose AI progress, ranked and analyzed.

Learning to Solve Hard Problems in RL for LLMs by Never Giving Up

· HF Daily Papers ·
The paper says RL training tends to reward problems models can already solve, leaving hard cases undertrained.

The authors call this a Matthew Effect: easy prompts get larger gains while hard prompts see smaller ones. Their proposed method, Never Give Up, keeps sampling on a problem until it gets a correct answer, shifting compute away from easy cases. On Deepscaler, they report better performance per compute, especially on harder math problems. On Manufactoria coding tasks, they say NGU keeps improving through harder tests where standard GRPO stalls. HF Daily Papers' note

score 5

Categories: Research