newsReddit r/LocalLLaMATrust 52 · CommunityPublished 10d agoLive · 9d ago
deepseek-v4-flash-0731 - surprisingly usable
I just finished building my (relatively) low rent local inference machine: * Epyc 7663 * 256GB ECC DDR4-3200 * 1x RTX 5090 32GB Yeah I realize it's weird to throw a 5090 and 256GB of anything together and call it low end, but relative to ~151GB of weights it is. I'm running UD-Q8_K_XL and getting 23.8-24.6 tokens/sec, with pp ranging from 60 on the first prompt to 385 near the last (no doubt lots of caching) on tasks using 100-128k total context. I
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD →
- PossiblePossibly related (embedding) · 52%pythongiant/KVBoost →
- PossiblePossibly related (embedding) · 52%Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference →
- PossiblePossibly related (embedding) · 59%yanun0323/deepseek_ssd →
- PossiblePossibly related (embedding) · 54%PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization →
- PossiblePossibly related (embedding) · 61%yanun0323/Whallm →
Covers
paperSLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPODrepopythongiant/KVBoostpaperDaedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inferencerepoyanun0323/deepseek_ssdpaperPagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization
Covers (incoming)
Related across the graph
repoyanun0323/deepseek_ssdrepoyanun0323/WhallmpaperDaedalus-150M: A Convolution-Attention Hybrid Designed for CPU InferencepaperSLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPODpaperPagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantizationrepopythongiant/KVBoost
