repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
helasaoudi/llm-inspector
The htop for LLM inference see exactly where every GB of VRAM goes and get measured quantization savings.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 56%Getting close to 100K context on 32GB VRAM with Qwen3.6-27 at Q8 →
- PossiblePossibly related (embedding) · 55%Going from single GPU to dual GPU is nice but not in the way I expected →
- PossiblePossibly related (embedding) · 53%I mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset) →
- PossiblePossibly related (embedding) · 52%Biggest, baddest model to fill 144GB VRAM + 120GB RAM to the brim, regardless of speed →
- PossiblePossibly related (embedding) · 45%Linux Improves VRAM Management in 7.3 Kernel 🥳 →
- PossiblePossibly related (embedding) · 48%Hybrid HBM-HBF Architecture in LLM Inference (University of Oxford) - Semiconductor Engineering →
Covers
newsGetting close to 100K context on 32GB VRAM with Qwen3.6-27 at Q8newsGoing from single GPU to dual GPU is nice but not in the way I expectednewsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)newsBiggest, baddest model to fill 144GB VRAM + 120GB RAM to the brim, regardless of speed
Covers (incoming)
Related across the graph
newsHybrid HBM-HBF Architecture in LLM Inference (University of Oxford) - Semiconductor EngineeringnewsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)newsLinux Improves VRAM Management in 7.3 Kernel 🥳newsGetting close to 100K context on 32GB VRAM with Qwen3.6-27 at Q8newsBiggest, baddest model to fill 144GB VRAM + 120GB RAM to the brim, regardless of speednewsGoing from single GPU to dual GPU is nice but not in the way I expected
