DataFlex-RL: An Evaluation Platform for RLVR Data Policies
Uniform GRPO beat the untrained model, but the tested RLVR data-policy tweaks did not reliably beat uniform sampling.
The paper introduces DataFlex-RL as a platform for comparing rollout selection, reweighting, and adaptive mixture policies under a shared GRPO setup. In the main Qwen2.5-7B-Base experiment, uniform GRPO raised domain-balanced average accuracy by 7.76 percentage points across 12 benchmarks. Eight rollout-selection or reweighting methods and three adaptive mixtures failed to show paired 95% confidence intervals excluding zero against their uniform or fixed-mixture baselines. The authors also report that a math-heavy six-benchmark summary reversed rankings relative to the 12-benchmark domain-balanced summary, underscoring how sensitive these evaluations are to benchmark choice. HF Daily Papers' note
The paper introduces DataFlex-RL as a platform for comparing rollout selection, reweighting, and adaptive mixture policies under a shared GRPO setup. In the main Qwen2.5-7B-Base experiment, uniform GRPO raised domain-balanced average accuracy by 7.76 percentage points across 12 benchmarks. Eight rollout-selection or reweighting methods and three adaptive mixtures failed to show paired 95% confidence intervals excluding zero against their uniform or fixed-mixture baselines. The authors also report that a math-heavy six-benchmark summary reversed rankings relative to the 12-benchmark domain-balanced summary, underscoring how sensitive these evaluations are to benchmark choice. HF Daily Papers' note
score 5