newsReddit r/MachineLearningTrust 72 · CommunityPublished 1mo agoLive · 1mo ago
Adaptive Mixture of Experts Gate (AMG) [R]
[Project] Post-hoc Adaptive MoE Gating on Qwen3.6-35B — empirical benchmarking of an open research gap Adaptive MoE routing — selecting a variable number of experts per token based on routing confidence — has been studied in papers (XMoE 2024, DynMoE ICLR 2025, TopP routing Huang et al. 2024). All successful implementations train from scratch. Nobody has published empirical results for post-hoc application to a pretrained fixed-k model at
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownScaling laws for mixture-of-experts models →
- PossiblePossibly related (embedding) · 49%It Takes a MAESTRO To Prune Bad Experts →
- PossiblePossibly related (embedding) · 48%microsoft/Tutel →
