Read original ↗
newsReddit r/LocalLLaMATrust 52 · CommunityPublished 1mo agoLive · 1mo ago

DeepSeek V4 Flash | IQ3_XXS-AS & IQ2_S Bench | mainline b10064 vs fairydreaming | 1xRTX 3090 + 128GB DDR4 | 250PP/11TG on 50K CTX

Hey all! Wanted to see how DeepSeek V4 Flash GGUFs in two different quants perform on my hardware and share the results. Tested two quants on the fairydreaming/llama.cpp dsv4 fork . As a bonus, I also ran the same model (IQ3_XXS-AS) on mainline llama.cpp b10064 just to check. TLDR: turns out the fork is pointless now! Model DeepSeek V4 Flash made by

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Related across the graph