Read original ↗
paperarXivTrust 82 · PrimaryPublished 15d agoLive · 12d ago

ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs

Mixture-of-Experts (MoE) architectures provide an efficient paradigm for scaling large language models (LLMs), yet fixed top-k routing activates the same number of expert slots for every token, causing substantial redundant computation. Existing expert-skipping methods often rely on router confidence, calibration data, or additional training, and therefore cannot reliably estimate the actual contribution of routed experts. To this end, we propose ACE, a training-free, calibration-free, and checkpoint-preserving framework for token-adaptive expert skipping in MoE-based LLMs. ACE contains two co

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 64%Adaptive Mixture of Experts Gate (AMG) [R]
  • FuzzyOverlapping authors or contributors · 62%affaan-m/ECC

    Shared author/contributor keys: jiang

  • FuzzyOverlapping authors or contributors · 62%BerriAI/litellm

    Shared author/contributor keys: jiang

  • LinkedLinked via arxiv author · 85%Zukang Xu

    ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs

  • LinkedLinked via arxiv author · 85%Zhixiong Zhao

    ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs

  • LinkedLinked via arxiv author · 85%Xing Hu

    ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs

  • LinkedLinked via arxiv author · 85%Jiangyong Yu

    ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs

  • LinkedLinked via arxiv author · 85%Houji Wen

    ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs

Covers

Implements (incoming)

authored (incoming)

Related across the graph

Topics