Strategically Diverse Sampling for Self-Training
The paper says diverse reasoning strategies can beat bigger-teacher distillation, even when the traces are wrong.
The authors test self-training data selected for “strategic diversity,” meaning different approaches to the same problem, instead of IID samples filtered mainly for correctness. They introduce GROOT, which samples from a hierarchy of approaches, and adapt Verbalized Sampling for unstructured approach sets. On competitive programming and Next-Chapter Prediction tasks, strategically sampled data improves difficult-task performance and helps later RL and test-time scaling. The strongest claim is that Qwen3-4B self-trained on diverse incorrect traces outperforms IID distillation from a 235B teacher. ArXiv · AI/CL/LG's note
The authors test self-training data selected for “strategic diversity,” meaning different approaches to the same problem, instead of IID samples filtered mainly for correctness. They introduce GROOT, which samples from a hierarchy of approaches, and adapt Verbalized Sampling for unstructured approach sets. On competitive programming and Next-Chapter Prediction tasks, strategically sampled data improves difficult-task performance and helps later RL and test-time scaling. The strongest claim is that Qwen3-4B self-trained on diverse incorrect traces outperforms IID distillation from a 235B teacher. ArXiv · AI/CL/LG's note
score 5