repoGitHubTrust 82 · PrimaryPublished 2d agoLive · 2d ago
helasaoudi/llm-inspector
The htop for LLM inference see exactly where every GB of VRAM goes and get measured quantization savings.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 56%Getting close to 100K context on 32GB VRAM with Qwen3.6-27 at Q8 →
- PossiblePossibly related (embedding) · 55%Going from single GPU to dual GPU is nice but not in the way I expected →
- PossiblePossibly related (embedding) · 53%I mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset) →
- PossiblePossibly related (embedding) · 52%Biggest, baddest model to fill 144GB VRAM + 120GB RAM to the brim, regardless of speed →
Covers
newsGetting close to 100K context on 32GB VRAM with Qwen3.6-27 at Q8newsGoing from single GPU to dual GPU is nice but not in the way I expectednewsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)newsBiggest, baddest model to fill 144GB VRAM + 120GB RAM to the brim, regardless of speed
Related across the graph
newsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)newsGetting close to 100K context on 32GB VRAM with Qwen3.6-27 at Q8newsBiggest, baddest model to fill 144GB VRAM + 120GB RAM to the brim, regardless of speednewsGoing from single GPU to dual GPU is nice but not in the way I expected
