newsReddit r/LocalLLaMATrust 52 · CommunityPublished 1mo agoLive · 1mo ago
GLM-5.2 on 8xB200: the deployment math nobody spells out - NVFP4 + 2x TP=4 replicas should beat TP=8 by ~2x. Full config guidance inside.
We have 8xB200 nodes and users keep asking us how to serve GLM-5.2 on them. Our engineering team went through everything published so far, and the optimal config is not the obvious one. Sharing the analysis because most of it applies wherever you rent or rack your B200s. The model GLM-5.2: ~750B total / ~40B active MoE (256 experts, top-8 routing, ~5.9% sparsity), DSA + MLA attention, 1M context, MIT license. Weights: ~744 GB in FP
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 47%SemiAnalysisAI/InferenceX →
