Megadose Built for builders and researchers.

Decoupling Exploration from Optimization in RLVR

· ArXiv · AI/CL/LG ·
ExpDis separates novelty-seeking rollouts from the policy that ultimately gets optimized.

The paper trains explorer policies with a novelty bonus, filters their outputs for correctness and quality, then distills those trajectories into a student policy trained without that bonus. The authors say this avoids the quality drops seen when novelty incentives are applied directly in RLVR. Across seven math reasoning benchmarks and two model families, ExpDis beat DAPO under the same wall-clock budget. They also report better pass@k scaling, suggesting more diverse correct solutions.

ArXiv · AI/CL/LG's note

score 5

Categories: Research