paperarXivTrust 82 · PrimaryPublished 3mo agoLive · 3mo ago
Scaling laws for mixture-of-experts models
How sparse expert routing changes the compute-optimal frontier for large models.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownIntroducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains →
- LinkedLinked via unknownAdaptive Mixture of Experts Gate (AMG) [R] →
