Megadose AI progress, ranked and analyzed.

PaperGym: Rubric-Centered Evolution for Research-Plan Generation

· ArXiv · AI/CL/LG ·
PaperGym turns research papers into training environments for AI research-plan generation.

The paper says research plans are hard to train with reinforcement learning because they lack verifiable answers and task-specific critics. Its framework synthesizes questions from a paper’s goal and background, then derives evaluation criteria from the method and experiments to reduce reward-by-paraphrase. The authors report lower criterion leakage than existing datasets and stronger results across Qwen3 model sizes than supervised fine-tuning or either training stage alone. They also release the PaperGym pipeline, a 20,000-instance corpus, benchmarks, and a trained model. ArXiv · AI/CL/LG's note

score 5

Categories: Research