newsReddit r/LocalLLaMATrust 52 · CommunityPublished 5d agoLive · 5d ago
Qwen3.8-Flash on RTX3090 + 64GB RAM (but you only need 12GB VRAM)
I've got Qwen3.8-Flash-next running on RTX 3090, Ryzen 9 3950X, a PCIe 3.0 motherboard, and 64GB DDR RAM from 2020. IQ4_XS weights, full kvarn5 context, vision on GPU, experts in host RAM, n-grams on disk. MTP works but actually slows decode down even with 80% draft acceptance, as expected since every rejected token eats into the host RAM bandwidth. I get 160 tok/s prefill 16 tok/s decode , which makes it a decent option whene
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%yanun0323/Whallm →
- PossiblePossibly related (embedding) · 49%AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-DFlash →
- PossiblePossibly related (embedding) · 63%carloslfu/slotstream →
