Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining
We uncover ELR collapse in language model pretraining: learning rate (LR) and parameter norm govern loss dynamics primarily through their ratio, the effective learning rate (ELR). When ELR is matched across runs, their loss trajectories collapse throughout training despite substantially different LRs and parameter norms. Across optimizers, architectures, datasets, and model scales, mean collapse errors are typically a few x 10^-3, below the seed-to-seed variation measured in a representative configuration. Systematic ablations identify normalization design and the timescale of LR-norm variatio
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 47%Dispersion loss counteracts embedding condensation in small language models →
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzySimilar title/name (fuzzy) · 59%aymericdamien/TopDeepLearning →
“Fuzzy title match (0.73): “Effective Learning Rate Governs Loss Dynamics in Language Mo” ≈ “aymericdamien/TopDeepLearning””
- LinkedLinked via arxiv author · 85%Zihan Liu →
“Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining”
- LinkedLinked via arxiv author · 85%Ruiheng Zheng →
“Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining”
- LinkedLinked via arxiv author · 85%Shaobo Zhang →
“Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining”
- LinkedLinked via arxiv author · 85%Changxin Tian →
“Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining”
- LinkedLinked via arxiv author · 85%Kunlong Chen →
“Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining”
