newsReddit r/LocalLLaMATrust 58 · CommunityPublished 1mo agoLive · 1mo ago
Devs - you have 64gb of VRAM - which model do you use for coding?
I've currently settled on an unsloth version of Qwen 3.5 122b-a10b model (UD-IQ4_NL). With 100k bf16 context window, I only had to load a few layers into CPU/RAM, it runs around 30 tok/sec which is fine for me. I've tested many models, hours of testing but I am currently deeply impressed with this one. I also use the Qwen 3.6 models (both) depending on need, but I think this biggun' is about to become my daily driver. Curious to know what others wi
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 47%LMCache/LMCache →
- PossiblePossibly related (embedding) · 47%mlhher/late-cli →
- PossiblePossibly related (embedding) · 51%Long-Context Fine-Tuning with Limited VRAM →
- PossiblePossibly related (embedding) · 50%Helldez/BigMoeOnEdge →
- PossiblePossibly related (embedding) · 50%gqgs/llm100kbench →
