TuringLLM: Efficiently Scaling Foundation Models Toward Physical AI
We present Turing-20B-A2B, a 20B-parameter Mixture-of-Experts language model that activates approximately 2B parameters per token, designed for long-context and latency-sensitive physical AI applications. The model adopts Quantile Routing in a dynamic top-k configuration, enabling token-adaptive expert allocation while maintaining balanced expert utilization and a controlled average compute budget. During deployment, we further apply capacity-constrained routing to prompt prefill for more regular and efficient expert execution, while retaining dropless routing during pretraining. Turing-20B-A2
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%Adaptive Mixture of Experts Gate (AMG) [R] →
- PossiblePossibly related (embedding) · 49%AT&T and Microsoft scale trillion-token workloads with Microsoft Foundry and AMD - Microsoft Azure →
- PossiblePossibly related (embedding) · 48%Meta researchers taught an 8B AI model to match Claude Opus 4.5 — without the frontier price tag →
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%sgl-project/sglang →
“Shared author/contributor keys: zhou”
- LinkedLinked via arxiv author · 85%Yuheng Zhang →
“TuringLLM: Efficiently Scaling Foundation Models Toward Physical AI”
