Read original ↗
paperarXivTrust 82 · PrimaryPublished 4d agoLive · 22h ago

Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

We study the scaling behavior of learning rate and batch size in pretraining dense large language models on English-prevalent corpora. Beyond scaling \textit{jointly optimal} learning rates and batch sizes, we investigate their \textit{marginal} evolution with model capacity and data scale and develop a model that captures these relationships. As we employ a Warmup-Stable-Decay learning rate schedule, we further investigate the gains from learning rate annealing over a broad range of hyperparameters settings, models and data budgets, and whether the optimal learning rate and batch size \textit

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Niccolò Ajroldi

    Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

  • LinkedLinked via arxiv author · 85%Diana Alexandra Onutu

    Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

  • LinkedLinked via arxiv author · 85%Haider Al-Tahan

    Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

  • LinkedLinked via arxiv author · 85%Jörg Franke

    Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

  • LinkedLinked via arxiv author · 85%Sampo Pyysalo

    Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

  • LinkedLinked via arxiv author · 85%Jenia Jitsev

    Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

  • LinkedLinked via arxiv author · 85%Aaron Klein

    Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

authored (incoming)

Related across the graph

Topics