Curvature-Conditioned Multiscale Momentum with Sphere Constraints for LLM Pretraining
Pretraining accounts for a large fraction of the total computational cost in LLM training. However, noise-dominant gradients and the highly ill-conditioned loss landscape bring severe challenges. Although modern adaptive optimizers such as AdamW and Muon have achieved great success in large-scale pretraining, their reliance on gradient normalization offers limited mitigation of the ill-conditioned curvature. The progress along flat directions (eigen-directions of small eigenvalues), which dominates the final loss reduction, remains relatively slow. To enhance training dynamics along flat direc
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%H64LM: A 249M-parameter Mixture-of-Experts Transformer built from scratch in PyTorch [P] →
- PossiblePossibly related (embedding) · 46%TorchJD: Training with multiple losses in PyTorch [P] →
- LinkedLinked via arxiv author · 85%Shuchen Zhu →
“Curvature-Conditioned Multiscale Momentum with Sphere Constraints for LLM Pretraining”
- LinkedLinked via arxiv author · 85%Yuxin Fang →
“Curvature-Conditioned Multiscale Momentum with Sphere Constraints for LLM Pretraining”
- LinkedLinked via arxiv author · 85%Mingze Wang →
“Curvature-Conditioned Multiscale Momentum with Sphere Constraints for LLM Pretraining”
- LinkedLinked via arxiv author · 85%Kun Yuan →
“Curvature-Conditioned Multiscale Momentum with Sphere Constraints for LLM Pretraining”
