Decoupling Exploration from Optimization in RLVR
ExpDis separates novelty-seeking rollouts from the policy that ultimately gets optimized.
The paper trains explorer policies with a novelty bonus, filters their outputs for correctness and quality, then distills those trajectories into a student policy trained without that bonus. The authors say this avoids the quality drops seen when novelty incentives are applied directly in RLVR. Across seven math reasoning benchmarks and two model families, ExpDis beat DAPO under the same wall-clock budget. They also report better pass@k scaling, suggesting more diverse correct solutions.
ArXiv · AI/CL/LG's note
The paper trains explorer policies with a novelty bonus, filters their outputs for correctness and quality, then distills those trajectories into a student policy trained without that bonus. The authors say this avoids the quality drops seen when novelty incentives are applied directly in RLVR. Across seven math reasoning benchmarks and two model families, ExpDis beat DAPO under the same wall-clock budget. They also report better pass@k scaling, suggesting more diverse correct solutions.
ArXiv · AI/CL/LG's note
score 5