Finetuning with Sampling: SFT Learns Better Than You Think
The paper argues SFT can match or beat RL-style posttraining when expert traces are resampled toward the model’s own distribution.
The authors introduce an MCMC sampling method that progressively makes off-policy data more on-policy for a reference model before finetuning. In their reported tests across scientific skills, math reasoning, and open-ended expertise, SFT with this sampling rivaled prevailing posttraining methods. They say it often generalized better and forgot less than strong on-policy baselines. ArXiv · AI/CL/LG's note
The authors introduce an MCMC sampling method that progressively makes off-policy data more on-policy for a reference model before finetuning. In their reported tests across scientific skills, math reasoning, and open-ended expertise, SFT with this sampling rivaled prevailing posttraining methods. They say it often generalized better and forgot less than strong on-policy baselines. ArXiv · AI/CL/LG's note
score 5