newsReddit r/LocalLLaMATrust 52 · CommunityPublished 17d agoLive · 16d ago
Qwen 3.8 2.4T at 288k tokens/s on Nvidia GB300 NVL72
https://developer.nvidia.com/blog/serve-qwen3-8-2-4t-a95b-a-2-4t-parameter-model-with-configurable-reasoning-on-nvidia-gb300-nvl72/ 4k tokens per second per GPU of which there are 72. 350 tokens per second per user "Without additional model tuning, the model achieves a throughput of over 4K tokens pe
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 46%giannisanni/pulsar →
