Megadose AI progress, ranked and analyzed.

Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior

· HF Daily Papers ·
The paper says synthetic pre-pretraining still buys token efficiency at 7B-scale tests, but the benefit does not track grammar.

The authors test five synthetic PPT tasks across four data mixtures, model sizes from 500M to 7B, and pretraining budgets up to 100B tokens. They report that PPT gains persist at scale, including at least 21B saved pretraining tokens at the 3B scale. But they find no consistent link between those gains and grammatical acceptability. The useful signal instead comes from PPT tasks that improve long-range retrieval, and the gains weaken mainly when web text is removed from the training mix. HF Daily Papers' note

score 5

Categories: Research