newsHacker NewsPublished yesterdayLive · yesterday
Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
I built slotstream, a way to run Qwen3.8-Flash-Next 4-bit on a low-memory mac starting from 16GB, a 125B parameter model that would need 100GB+ memory/RAM, thanks to expert-offloading/ssd-streaming. Easy to install/update, and mac-native using MLX and Swift. It ships with auto-mode, which makes a good tradeoff between memory usage and speed. I'll be implementing and porting the MTP module for speculative decoding next Comments URL:
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 59%yanun0323/Whallm →
- PossiblePossibly related (embedding) · 54%quantumnic/ssd-llm →
- PossiblePossibly related (embedding) · 52%sgl-project/SpecForge →
- PossiblePossibly related (embedding) · 51%yanun0323/deepseek_ssd →
- PossiblePossibly related (embedding) · 73%carloslfu/slotstream →
