repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 25d ago
ModelEngine-Group/unified-cache-management
Persist and reuse KV Cache to speedup your LLM.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 57%I mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset) →
- PossiblePossibly related (embedding) · 52%Evaluate a model properly →
- PossiblePossibly related (embedding) · 48%I compiled LLM inference pricing across 7 providers — the caching numbers are surprising(spreadsheet included) [R] →
- PossiblePossibly related (embedding) · 46%Biggest, baddest model to fill 144GB VRAM + 120GB RAM to the brim, regardless of speed →
- PossiblePossibly related (embedding) · 46%Best tps can I get with Qwen3.5 122B on 32GB VRAM + 64GB RAM? →
- PossiblePossibly related (embedding) · 54%Llama-Server is Throwing Away Your Perfectly Good KV Caches, and How to Fix It →
- PossiblePossibly related (embedding) · 46%If you're building a harness, here is a simple tool to catch cache invalidation in your calls to LLMs →
Covers
newsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)newsI compiled LLM inference pricing across 7 providers — the caching numbers are surprising(spreadsheet included) [R]newsBiggest, baddest model to fill 144GB VRAM + 120GB RAM to the brim, regardless of speednewsBest tps can I get with Qwen3.5 122B on 32GB VRAM + 64GB RAM?
Related to
Covers (incoming)
Related across the graph
newsIf you're building a harness, here is a simple tool to catch cache invalidation in your calls to LLMsnewsI compiled LLM inference pricing across 7 providers — the caching numbers are surprising(spreadsheet included) [R]newsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)newsBest tps can I get with Qwen3.5 122B on 32GB VRAM + 64GB RAM?newsLlama-Server is Throwing Away Your Perfectly Good KV Caches, and How to Fix IttutorialEvaluate a model properlynewsBiggest, baddest model to fill 144GB VRAM + 120GB RAM to the brim, regardless of speed
