Read original ↗
newsReddit r/LocalLLaMATrust 52 · CommunityPublished 29d agoLive · 28d ago

DeepSeek v4 Flash on 5090 in llama.cpp with 1 Million context

After the recent llama.cpp changes, DeepSeek V4 Flash has become much more usable. I ran some benchmarks and wanted to share the results along with the config I used. I'm using DeepSeek-V4-Flash-UD-Q8_K_XL from Unsloth: https://huggingface.co/unsloth/DeepSeek-V4-Flash-GGUF Config: llama-server \ -m

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Related across the graph