repoGitHubTrust 82 · PrimaryPublished 5d agoLive · 5d ago
yanun0323/Whallm
DeepSeek-V4-Flash-0731 284B inference in ~30 GB of RAM / Qwen3.8-Next-Flash-FP8 inference in ~20 GB of RAM on any M-series MacBook
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 61%deepseek-v4-flash-0731 - surprisingly usable →
- PossiblePossibly related (embedding) · 61%Mac Studio M5 Max Cost Analysis →
- PossiblePossibly related (embedding) · 59%DeepSeek V4 Flash (98GB) on 1x 4060ti + CPU got 300% faster this week [ 2->7t/s] →
- PossiblePossibly related (embedding) · 58%DeepSeek-V4-Flash (MXFP4): compute buffer scales ~3x just from KV cache quant type (f16 vs q8_0) — anyone else seeing this? Llama.cpp →
- PossiblePossibly related (embedding) · 57%Running DeepSeek V4 Flash Q4_K_XL at ~100 tok/s prompt processing on 4× RTX 3060 12GB →
- PossiblePossibly related (embedding) · 50%Framework Desktop Is Getting A 192GB RAM Boost For Monster Local LLMs - HotHardware →
- PossiblePossibly related (embedding) · 55%Qwen3.8-Flash-Next in llama.cpp from CPU-only to 96GB VRAM: 8.5 to 109 tok/s, max context and parameters test. My findings on RTX 6000 PRO. →
- PossiblePossibly related (embedding) · 46%Qwen 3.8 27b (Q4KM) oneshot a Super Mario clone →
Covers
newsdeepseek-v4-flash-0731 - surprisingly usablenewsMac Studio M5 Max Cost AnalysisnewsDeepSeek V4 Flash (98GB) on 1x 4060ti + CPU got 300% faster this week [ 2->7t/s]newsDeepSeek-V4-Flash (MXFP4): compute buffer scales ~3x just from KV cache quant type (f16 vs q8_0) — anyone else seeing this? Llama.cppnewsRunning DeepSeek V4 Flash Q4_K_XL at ~100 tok/s prompt processing on 4× RTX 3060 12GB
Covers (incoming)
newsFramework Desktop Is Getting A 192GB RAM Boost For Monster Local LLMs - HotHardwarenewsQwen3.8-Flash-Next in llama.cpp from CPU-only to 96GB VRAM: 8.5 to 109 tok/s, max context and parameters test. My findings on RTX 6000 PRO.newsQwen 3.8 27b (Q4KM) oneshot a Super Mario clonenewsShow HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/snewsQwen3.8-Flash on RTX3090 + 64GB RAM (but you only need 12GB VRAM)newsMac mini M6 16GB vs 24GB vs 32GB: Which Memory Should You Buy? - zeera wirelessnewsExo labs claiming 4.8 tb/s memory bandwidth through m5u Mac Studio clusteringnewsExperience report - Qwen 3.8 Flash Next on memory rich, GPU poor setup
Related across the graph
newsRunning DeepSeek V4 Flash Q4_K_XL at ~100 tok/s prompt processing on 4× RTX 3060 12GBnewsMac mini M6 16GB vs 24GB vs 32GB: Which Memory Should You Buy? - zeera wirelessnewsDeepSeek-V4-Flash (MXFP4): compute buffer scales ~3x just from KV cache quant type (f16 vs q8_0) — anyone else seeing this? Llama.cppnewsDeepSeek V4 Flash (98GB) on 1x 4060ti + CPU got 300% faster this week [ 2->7t/s]newsMac Studio M5 Max Cost Analysisnewsdeepseek-v4-flash-0731 - surprisingly usablenewsFramework Desktop Is Getting A 192GB RAM Boost For Monster Local LLMs - HotHardwarenewsQwen 3.8 27b (Q4KM) oneshot a Super Mario clonenewsQwen3.8-Flash-Next in llama.cpp from CPU-only to 96GB VRAM: 8.5 to 109 tok/s, max context and parameters test. My findings on RTX 6000 PRO.newsExperience report - Qwen 3.8 Flash Next on memory rich, GPU poor setupnewsQwen3.8-Flash on RTX3090 + 64GB RAM (but you only need 12GB VRAM)newsExo labs claiming 4.8 tb/s memory bandwidth through m5u Mac Studio clusteringnewsShow HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
