Megadose AI progress, ranked and analyzed.

Scaling Domain Data Repetition in LLM Pretraining

· HF Daily Papers ·
The paper finds that optimal domain-data repetition can rise with model size when tokens per parameter stay fixed.

The authors study repeated use of scarce, high-quality domain data as pretraining budgets scale with larger LLMs. They report that domains with lower final validation loss tend to tolerate or benefit from more repetition, while the amount of unique domain data is only weakly tied to the best repeat count. The result points to smaller proxy models, matched on tokens per parameter, as a practical way to estimate repetition settings for larger runs. HF Daily Papers' note

score 5

Categories: Research