Megadose AI progress, ranked and analyzed.

Understanding Reasoning from Pretraining to Post-Training

· ArXiv · AI/CL/LG ·
The paper argues that RL gains are strongly shaped by what pretraining already produced.

Using chess as a controlled testbed, the authors train models from 5M to 1B parameters through pretraining, supervised fine-tuning, and RL. They find post-RL performance at a given compute level can be predicted from pretraining loss, and that more pretraining tokens improve RL reward-curve slopes roughly linearly. RL behaves differently by difficulty: it reinforces already-likely correct moves on easy puzzles, but can surface correct moves that were nearly absent after SFT on hard ones. A smaller math-domain test shows the same pattern, with longer-pretrained checkpoints improving faster under RL. ArXiv · AI/CL/LG's note

score 5

Categories: Research