repoGitHubTrust 82 · PrimaryPublished yesterdayLive · yesterday
carloslfu/slotstream
Run Qwen3.8-Flash-Next (125B MoE, 104 GB at 4-bit) on Macs with a fraction of that RAM by streaming experts from SSD. MLX + Swift, Ollama-compatible API.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 73%Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s →
- PossiblePossibly related (embedding) · 63%Qwen3.8-Flash on RTX3090 + 64GB RAM (but you only need 12GB VRAM) →
- PossiblePossibly related (embedding) · 57%Mac Studio M5 Max Cost Analysis →
- PossiblePossibly related (embedding) · 57%Devs - you have 64gb of VRAM - which model do you use for coding? →
- PossiblePossibly related (embedding) · 54%It's unbelievable! I used the mmap function in llama.cpp to fit Qwen3.8-Flash-Next IQ3_XSS into 16G+64G RAM, and the speed still reached 26t/s. →
Covers
newsShow HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/snewsQwen3.8-Flash on RTX3090 + 64GB RAM (but you only need 12GB VRAM)newsMac Studio M5 Max Cost AnalysisnewsDevs - you have 64gb of VRAM - which model do you use for coding?newsIt's unbelievable! I used the mmap function in llama.cpp to fit Qwen3.8-Flash-Next IQ3_XSS into 16G+64G RAM, and the speed still reached 26t/s.
Related across the graph
newsIt's unbelievable! I used the mmap function in llama.cpp to fit Qwen3.8-Flash-Next IQ3_XSS into 16G+64G RAM, and the speed still reached 26t/s.newsMac Studio M5 Max Cost AnalysisnewsDevs - you have 64gb of VRAM - which model do you use for coding?newsQwen3.8-Flash on RTX3090 + 64GB RAM (but you only need 12GB VRAM)newsShow HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
