Read original ↗
newsReddit r/LocalLLaMATrust 52 · CommunityPublished 1mo agoLive · 1mo ago

GLM-5.2 on 8xB200: the deployment math nobody spells out - NVFP4 + 2x TP=4 replicas should beat TP=8 by ~2x. Full config guidance inside.

We have 8xB200 nodes and users keep asking us how to serve GLM-5.2 on them. Our engineering team went through everything published so far, and the optimal config is not the obvious one. Sharing the analysis because most of it applies wherever you rent or rack your B200s. The model GLM-5.2: ~750B total / ~40B active MoE (256 experts, top-8 routing, ~5.9% sparsity), DSA + MLA attention, 1M context, MIT license. Weights: ~744 GB in FP

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Related across the graph