repoGitHubTrust 82 · PrimaryPublished 22d agoLive · 1h ago
FedericoTs/quantprobe
Run a 110B on a 2016 PC with 16 GB RAM. Know your tok/s before you download. Placement beats budget: predicts speed + memory fit for any GGUF on your exact hardware, self-calibrates, emits the exact llama.cpp command — or 'quantprobe auto' does it all. Falsification-tested laws; misses published at full size. pip install quantprobe
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 68%I mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset) →
- PossiblePossibly related (embedding) · 60%Best Local LLMs You Can Run on a Single 24GB GPU in 2026: Qwen, Gemma, Mistral, DeepSeek Compared - MarkTechPost →
- PossiblePossibly related (embedding) · 59%OpenAI and Broadcom announce chip designed for LLM inference at scale →
- PossiblePossibly related (embedding) · 57%I feel like I'm not using my hardware efficiently →
- PossiblePossibly related (embedding) · 57%GLM 5.2 running on MacBook Pro M5 48 GB Ram at between 2 - 2.8t/s →
- PossiblePossibly related (embedding) · 48%GLM 5.2 and ik_llama.ccp →
Covers
newsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)newsBest Local LLMs You Can Run on a Single 24GB GPU in 2026: Qwen, Gemma, Mistral, DeepSeek Compared - MarkTechPostnewsOpenAI and Broadcom announce chip designed for LLM inference at scalenewsI feel like I'm not using my hardware efficientlynewsGLM 5.2 running on MacBook Pro M5 48 GB Ram at between 2 - 2.8t/s
Covers (incoming)
Related across the graph
newsOpenAI and Broadcom announce chip designed for LLM inference at scalenewsGLM 5.2 running on MacBook Pro M5 48 GB Ram at between 2 - 2.8t/snewsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)newsBest Local LLMs You Can Run on a Single 24GB GPU in 2026: Qwen, Gemma, Mistral, DeepSeek Compared - MarkTechPostnewsI feel like I'm not using my hardware efficientlynewsGLM 5.2 and ik_llama.ccp
