newsReddit r/LocalLLaMATrust 52 · CommunityPublished 7d agoLive · 7d ago
Mac Studio M5 Max Cost Analysis
At $10k, you could get - 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan) - 5.7B tokens with DeepSeek V4 Pro OpenRouter - 100B tokens with DeepSeek V4 Flash OpenRouter As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 59%waybarrios/vllm-mlx →
- PossiblePossibly related (embedding) · 54%marzukia/qMLX →
- PossiblePossibly related (embedding) · 54%raullenchai/Rapid-MLX →
- PossiblePossibly related (embedding) · 52%openvinotoolkit/model_server →
- PossiblePossibly related (embedding) · 52%yanun0323/deepseek_ssd →
- PossiblePossibly related (embedding) · 53%Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 →
- PossiblePossibly related (embedding) · 61%yanun0323/Whallm →
