repoGitHubTrust 82 · PrimaryPublished 3h agoLive · 3h ago
yanun0323/deepseek_ssd
DeepSeek-V4-Flash-0731 284B inference in ~30 GB of RAM on any M-series MacBook
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%DeepSeek V4 Flash (98GB) on 1x 4060ti + CPU got 300% faster this week [ 2->7t/s] →
- PossiblePossibly related (embedding) · 51%Running DeepSeek V4 Flash Q4_K_XL at ~100 tok/s prompt processing on 4× RTX 3060 12GB →
- PossiblePossibly related (embedding) · 50%DeepSeek-V4-Flash (MXFP4): compute buffer scales ~3x just from KV cache quant type (f16 vs q8_0) — anyone else seeing this? Llama.cpp →
- PossiblePossibly related (embedding) · 52%I mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset) →
- PossiblePossibly related (embedding) · 52%DeepSeek v4 Flash on 4090 + DDR5, my experience →
Covers
newsDeepSeek V4 Flash (98GB) on 1x 4060ti + CPU got 300% faster this week [ 2->7t/s]newsRunning DeepSeek V4 Flash Q4_K_XL at ~100 tok/s prompt processing on 4× RTX 3060 12GBnewsDeepSeek-V4-Flash (MXFP4): compute buffer scales ~3x just from KV cache quant type (f16 vs q8_0) — anyone else seeing this? Llama.cppnewsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)newsDeepSeek v4 Flash on 4090 + DDR5, my experience
Related across the graph
newsRunning DeepSeek V4 Flash Q4_K_XL at ~100 tok/s prompt processing on 4× RTX 3060 12GBnewsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)newsDeepSeek-V4-Flash (MXFP4): compute buffer scales ~3x just from KV cache quant type (f16 vs q8_0) — anyone else seeing this? Llama.cppnewsDeepSeek V4 Flash (98GB) on 1x 4060ti + CPU got 300% faster this week [ 2->7t/s]newsDeepSeek v4 Flash on 4090 + DDR5, my experience
