Megadose AI progress, ranked and analyzed.

Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack

· ArXiv · AI/CL/LG ·
A stronger pretraining checkpoint did not necessarily make the best post-SFT model.

The paper says this failure showed up in a full 30B mixture-of-experts training pipeline. The authors found that checkpoints which held up better after downstream training also had higher “solution density,” meaning their performance survived local weight perturbations. The claim is narrower than benchmark skepticism: pretraining loss or benchmark scores alone can pick the wrong starting point for later training. ArXiv · AI/CL/LG's note

score 5

Categories: Research