newsReddit r/LocalLLaMATrust 58 · CommunityPublished 1mo agoLive · 1mo ago
DeepSeek-V4-Flash (MXFP4): compute buffer scales ~3x just from KV cache quant type (f16 vs q8_0) — anyone else seeing this? Llama.cpp
Bartowski's DeepSeek-V4-Flash-MXFP4 GGUF, llama.cpp build 9851 ( 0eca4d490 ), deepseek4 arch. Ran the same n_ctx = 10240 , same n_ubatch = n_batch = 8192 , flash attention on — only difference is -ctk / -ctv : Cache type Total KV cache (CUDA0) CUDA0 compute buffer
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 46%Zefan-Cai/R-KV →
- PossiblePossibly related (embedding) · 51%Menfre01/waveloom →
- PossiblePossibly related (embedding) · 52%xcena-dev/maru →
- PossiblePossibly related (embedding) · 53%test5630352/llm-cache-optimize →
- PossiblePossibly related (embedding) · 46%DaoyuanLi2816/mini-verl →
