newsReddit r/LocalLLaMATrust 58 · CommunityPublished 1mo agoLive · 1mo ago
Biggest, baddest model to fill 144GB VRAM + 120GB RAM to the brim, regardless of speed
I'm trying to round out my quiver of daily driver models for my personal harness. Right now I drive qwen3.6 27b for balanced code and gemma4 31b for human interaction with lots of context and a few parallel sessions. Minimax M2.7 at Q6 clocks in at 207gb base and just barely fits once I get KV cache and context down for when I have a "take all day to answer; just be right" problem. I'm debating on moving to M3 at Q3, but I'm wondering if there are any
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%LMCache/LMCache →
- PossiblePossibly related (embedding) · 47%mlhher/late-cli →
- PossiblePossibly related (embedding) · 46%ModelEngine-Group/unified-cache-management →
- PossiblePossibly related (embedding) · 55%Long-Context Fine-Tuning with Limited VRAM →
- PossiblePossibly related (embedding) · 52%helasaoudi/llm-inspector →
