newsReddit r/LocalLLaMATrust 58 · CommunityPublished 1mo agoLive · 1mo ago
I mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)
I kept answering the same question for friends ("I've got a 16GB MacBook / a 3060, what can I actually run?") and got tired of guessing, so I started a spreadsheet. It grew into a real dataset, so I put it on GitHub under CC BY for anyone to use or fix. Rule of thumb I landed on: at Q4_K_M a model needs roughly 0.6GB of memory per billion params, and you want to size to about 70% of your RAM/VRAM so the OS, context and KV cache still have room.
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownEvaluate a model properly →
- LinkedLinked via unknownGSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache →
- PossiblePossibly related (embedding) · 50%NVIDIA-NeMo/Curator →
- PossiblePossibly related (embedding) · 53%jmaczan/tiny-vllm →
- PossiblePossibly related (embedding) · 58%LMCache/LMCache →
- PossiblePossibly related (embedding) · 48%mlhher/late-cli →
- PossiblePossibly related (embedding) · 49%luziyao1995/vllm →
- PossiblePossibly related (embedding) · 49%jundot/omlx →
Covers
Covers (incoming)
paperGSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV CacherepoNVIDIA-NeMo/Curatorrepojmaczan/tiny-vllmrepoLMCache/LMCacherepomlhher/late-clirepoluziyao1995/vllmrepojundot/omlxrepojohn-rocky/apple-silicon-llm-benchreponovitalabs/pegaflowrepoModelEngine-Group/unified-cache-managementrepomanjunathshiva/turboquant-mlxrepotrvon/yamsrepoAndyyyy64/whichllmrepoohdearquant/latticerepopythongiant/KVBoostrepoRyan-Adams57/model-fitrepodevelopment-and-operations/model-fitrepoopenlake-project/openlakerepoVectifyAI/ConDBrepoxcena-dev/marurepoalibaba/tair-kvcacherepotest5630352/llm-cache-optimizerepojagmarques/nexusquantpaperLong-Context Fine-Tuning with Limited VRAMrepoovg-project/kvcachedrepopubmethod/run-local-llrepoHelldez/BigMoeOnEdgepaperPagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantizationrepohelasaoudi/llm-inspectorrepoShadowLLM/shadow-peftrepoFedericoTs/quantprobe
Related across the graph
repoRyan-Adams57/model-fitrepotest5630352/llm-cache-optimizerepoohdearquant/latticerepoFedericoTs/quantprobepaperLong-Context Fine-Tuning with Limited VRAMreponovitalabs/pegaflowrepoVectifyAI/ConDBrepoShadowLLM/shadow-peftrepoopenlake-project/openlakerepodevelopment-and-operations/model-fitrepojmaczan/tiny-vllmrepoAndyyyy64/whichllmrepoluziyao1995/vllmrepomlhher/late-clipaperGSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cacherepojundot/omlxrepoalibaba/tair-kvcacherepoLMCache/LMCacherepoModelEngine-Group/unified-cache-managementrepomanjunathshiva/turboquant-mlxtutorialEvaluate a model properlyrepopubmethod/run-local-llrepohelasaoudi/llm-inspectorrepoHelldez/BigMoeOnEdgerepotrvon/yamsrepojohn-rocky/apple-silicon-llm-benchrepoxcena-dev/marupaperPagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantizationrepopythongiant/KVBoostrepojagmarques/nexusquantrepoovg-project/kvcachedrepoNVIDIA-NeMo/Curator
