repoGitHubTrust 82 · PrimaryPublished 2mo agoLive · 5d ago
InternLM/xtuner
A Next-Generation Training Engine Built for Ultra-Large MoE Models
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%Qwen3.6-27B UD Q3 with kv at q8 is quite amazing for simple proof of concepts →
- PossiblePossibly related (embedding) · 47%Meituan Open-Sources 1.6-Trillion-Parameter LongCat-2.0, First Model Fully Trained on 50,000 Chinese-Made AI Chips - finance.biggo.com →
- PossiblePossibly related (embedding) · 45%Tencent-HY3 is the real deal on 128GB! →
- PossiblePossibly related (embedding) · 50%Benchmarks: AntLing-3.0-flash a hybrid-reasoning MoE model built for production-scale agents. →
- PossiblePossibly related (embedding) · 49%Don't want to be this guy, but I need Qwen 3.8 35B A3B →
Covers
Covers (incoming)
newsBenchmarks: AntLing-3.0-flash a hybrid-reasoning MoE model built for production-scale agents.newsDon't want to be this guy, but I need Qwen 3.8 35B A3BnewsProposed architecture for inferencing sparse MOE models increasing Active parameters using layered + linear decay. Succinct reasoning without any model training or fine tune. [p]
Related across the graph
newsProposed architecture for inferencing sparse MOE models increasing Active parameters using layered + linear decay. Succinct reasoning without any model training or fine tune. [p]newsDon't want to be this guy, but I need Qwen 3.8 35B A3BnewsBenchmarks: AntLing-3.0-flash a hybrid-reasoning MoE model built for production-scale agents.newsMeituan Open-Sources 1.6-Trillion-Parameter LongCat-2.0, First Model Fully Trained on 50,000 Chinese-Made AI Chips - finance.biggo.comnewsTencent-HY3 is the real deal on 128GB!newsQwen3.6-27B UD Q3 with kv at q8 is quite amazing for simple proof of concepts
