Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss
The paper sets a first scaling baseline for future OpenEuroLLM pretraining runs.
The authors study how learning rate and batch size scale for dense LLMs trained on English-prevalent corpora. They test both jointly optimal settings and how each hyperparameter changes with model size and data scale. They also examine gains from learning-rate annealing under a Warmup-Stable-Decay schedule, including whether settings transfer between stable and decay phases. Their loss-scaling analysis finds newer forms useful for modeling both undertraining and overtraining regimes. ArXiv · AI/CL/LG's note
The authors study how learning rate and batch size scale for dense LLMs trained on English-prevalent corpora. They test both jointly optimal settings and how each hyperparameter changes with model size and data scale. They also examine gains from learning-rate annealing under a Warmup-Stable-Decay schedule, including whether settings transfer between stable and decay phases. Their loss-scaling analysis finds newer forms useful for modeling both undertraining and overtraining regimes. ArXiv · AI/CL/LG's note
score 5