Read original ↗
newsReddit r/LocalLLaMATrust 58 · CommunityPublished 1mo agoLive · 1mo ago

Tesla V100 16GB local LLMs, single and dual NVLink benchmarks

Picked up a couple of Tesla V100-SXM2-16GB modules a while back to run local models and drive Claude Code fully offline, figured the actual numbers and the traps might save someone else the pain. They've come right down in price and the 16GB of HBM2 at ~900 GB/s still holds up surprisingly well for inference, bandwidth is what matters most for token gen and the V100 has heaps of it. Spec refresher: GV100, Volta, sm_70, 16GB HBM2 ~900 GB/s, fp16 on

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers (incoming)

Related across the graph