repoGitHubTrust 82 · PrimaryPublished 27d agoLive · 22d ago
Helldez/BigMoeOnEdge
Run a 120B-parameter MoE (60 GB) on a 12 GB phone. CPU-only, lossless, on stock llama.cpp
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%Reducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading - NVIDIA Developer →
- PossiblePossibly related (embedding) · 53%I feel like I'm not using my hardware efficiently →
- PossiblePossibly related (embedding) · 52%DeepSeek v4 Flash on 4090 + DDR5, my experience →
- PossiblePossibly related (embedding) · 51%Tried testing qwen 35b moe model on s26 ultra , without compromising on precision [R] ,[D] →
- PossiblePossibly related (embedding) · 50%I mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset) →
- PossiblePossibly related (embedding) · 51%Best tps can I get with Qwen3.5 122B on 32GB VRAM + 64GB RAM? →
- PossiblePossibly related (embedding) · 50%Devs - you have 64gb of VRAM - which model do you use for coding? →
- PossiblePossibly related (embedding) · 54%Best Local LLMs You Can Run on a Single 24GB GPU in 2026: Qwen, Gemma, Mistral, DeepSeek Compared - MarkTechPost →
Covers
newsReducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading - NVIDIA DevelopernewsI feel like I'm not using my hardware efficientlynewsDeepSeek v4 Flash on 4090 + DDR5, my experiencenewsTried testing qwen 35b moe model on s26 ultra , without compromising on precision [R] ,[D]newsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)newsBest tps can I get with Qwen3.5 122B on 32GB VRAM + 64GB RAM?newsDevs - you have 64gb of VRAM - which model do you use for coding?
Covers (incoming)
Related across the graph
newsM2 Ultra 64gb vs m1 ultra 128gbnewsReducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading - NVIDIA DevelopernewsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)newsTried testing qwen 35b moe model on s26 ultra , without compromising on precision [R] ,[D]newsBest Local LLMs You Can Run on a Single 24GB GPU in 2026: Qwen, Gemma, Mistral, DeepSeek Compared - MarkTechPostnewsDevs - you have 64gb of VRAM - which model do you use for coding?newsBest tps can I get with Qwen3.5 122B on 32GB VRAM + 64GB RAM?newsMSI Pro Max Edge AI+ Mini PC Runs 120B Local AI Models With 128GB RAM - HotHardwarenewsI feel like I'm not using my hardware efficientlynewsDeepSeek v4 Flash on 4090 + DDR5, my experience
