newsGoogle News — Machine LearningTrust 62 · AggregatorPublished 17d agoLive · 16d ago
Scaling Laws for Mixture Pretraining Under Data Constraints - Apple Machine Learning Research
Scaling Laws for Mixture Pretraining Under Data Constraints Apple Machine Learning Research
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 47%Fine-tuning →
- PossiblePossibly related (embedding) · 47%Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts →
- PossiblePossibly related (embedding) · 46%Scaling laws for mixture-of-experts models →
- PossiblePossibly related (embedding) · 45%CausalMix: Data Mixture as Causal Inference for Language Model Training →
- PossiblePossibly related (embedding) · 45%HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures →
Covers
glossary_termFine-tuningpaperLet's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-ExpertspaperScaling laws for mixture-of-experts modelspaperCausalMix: Data Mixture as Causal Inference for Language Model TrainingpaperHERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures
Related across the graph
paperScaling laws for mixture-of-experts modelsglossary_termFine-tuningpaperCausalMix: Data Mixture as Causal Inference for Language Model TrainingpaperHERMES: A Multi-Granularity Labeling Substrate for Pre-training Data MixturespaperLet's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts
