PaperGym: Rubric-Centered Evolution for Research-Plan Generation
PaperGym turns papers into training environments where rubrics act as both teacher context and reward.
The authors use a paper’s goal and background to synthesize the question, then derive evaluation criteria from its methods and experiments. They report lower criterion leakage than existing datasets, at 3.7%. Across Qwen3 model sizes, their training schedule beats supervised fine-tuning and single-stage variants on five-benchmark averages. They also release the PaperGym pipeline, a 20,000-instance corpus, benchmarks, and trained models. HF Daily Papers' note
The authors use a paper’s goal and background to synthesize the question, then derive evaluation criteria from its methods and experiments. They report lower criterion leakage than existing datasets, at 3.7%. Across Qwen3 model sizes, their training schedule beats supervised fine-tuning and single-stage variants on five-benchmark averages. They also release the PaperGym pipeline, a 20,000-instance corpus, benchmarks, and trained models. HF Daily Papers' note
score 5