Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss
We study the scaling behavior of learning rate and batch size in pretraining dense large language models on English-prevalent corpora. Beyond scaling \textit{jointly optimal} learning rates and batch sizes, we investigate their \textit{marginal} evolution with model capacity and data scale and develop a model that captures these relationships. As we employ a Warmup-Stable-Decay learning rate schedule, we further investigate the gains from learning rate annealing over a broad range of hyperparameters settings, models and data budgets, and whether the optimal learning rate and batch size \textit
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Niccolò Ajroldi →
“Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss”
- LinkedLinked via arxiv author · 85%Diana Alexandra Onutu →
“Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss”
- LinkedLinked via arxiv author · 85%Haider Al-Tahan →
“Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss”
- LinkedLinked via arxiv author · 85%Jörg Franke →
“Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss”
- LinkedLinked via arxiv author · 85%Sampo Pyysalo →
“Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss”
- LinkedLinked via arxiv author · 85%Jenia Jitsev →
“Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss”
- LinkedLinked via arxiv author · 85%Aaron Klein →
“Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss”
