It Takes a MAESTRO To Prune Bad Experts
Sparsely-activated Mixture-of-Experts (MoE) language models achieve remarkable inference efficiency by activating only a small fraction of parameters per token, yet their full expert banks reside in memory at all times, creating a prohibitive deployment bottleneck. Existing structured pruning methods, largely designed for dense transformers, assess expert importance using locally derived heuristics that are blind to the interdependent nature of MoE routing. We introduce MAESTRO (Markov-chain Approximated Expert Sparsification via Transition-based ROuting), a structured pruning framework design
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%H64LM: A 249M-parameter Mixture-of-Experts Transformer built from scratch in PyTorch [P] →
- PossiblePossibly related (embedding) · 49%Transformer →
- PossiblePossibly related (embedding) · 49%Adaptive Mixture of Experts Gate (AMG) [R] →
- PossiblePossibly related (embedding) · 49%chrisliu298/awesome-llm-unlearning →
- PossiblePossibly related (embedding) · 46%thu-pacman/chitu →
- LinkedLinked via arxiv author · 85%Palaash Goel →
“It Takes a MAESTRO To Prune Bad Experts”
- LinkedLinked via arxiv author · 85%Ayush Maheshwari →
“It Takes a MAESTRO To Prune Bad Experts”
- LinkedLinked via arxiv author · 85%Tanmoy Chakraborty →
“It Takes a MAESTRO To Prune Bad Experts”
