newsReddit r/LocalLLaMATrust 52 · CommunityPublished 29d agoLive · 28d ago
DeepSeek v4 Flash on 5090 in llama.cpp with 1 Million context
After the recent llama.cpp changes, DeepSeek V4 Flash has become much more usable. I ran some benchmarks and wanted to share the results along with the config I used. I'm using DeepSeek-V4-Flash-UD-Q8_K_XL from Unsloth: https://huggingface.co/unsloth/DeepSeek-V4-Flash-GGUF Config: llama-server \ -m
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 48%usewhale/Whale →
- PossiblePossibly related (embedding) · 47%Eeveeboo/e2e-roleplay-testing →
- PossiblePossibly related (embedding) · 47%Menfre01/waveloom →
- PossiblePossibly related (embedding) · 46%deepseek-ai/DeepSeek-V4-Pro →
- PossiblePossibly related (embedding) · 46%luojieLLMaaS/haxiv →
