Megadose AI progress, ranked and analyzed.

Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space

· HF Daily Papers ·
RLVR’s lost reasoning diversity shows up before the first arithmetic move, not after it.

The paper tests Countdown tasks where possible solution “entrances” can be enumerated by first operand and operator. Across PPO and GRPO runs on Qwen2.5 models, solution coverage drops by up to 67% even as pass@1 improves. Likelihood shifts are 11x-16x larger before the first arithmetic operation than later in the reasoning path. When researchers provide an otherwise unchosen entrance prefix, low-access solution families become executable again, and late-layer interpolation with earlier checkpoints raises coverage without hurting pass@1. HF Daily Papers' note

score 5

Categories: Research