repoGitHubTrust 82 · PrimaryPublished 2mo agoLive · 2mo ago
pythongiant/KVBoost
Make local LLM inference faster with chunk-level KV cache reuse
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 56%I mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset) →
- PossiblePossibly related (embedding) · 51%Hardware startup unveils inference accelerator →
- PossiblePossibly related (embedding) · 49%Would having a dedicated programming language specifically for LLMs be a viable solution? [D] →
- PossiblePossibly related (embedding) · 48%RaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM Inference →
- PossiblePossibly related (embedding) · 48%Evaluate a model properly →
- PossiblePossibly related (embedding) · 45%I merged fixes for quantized KV cache into my DeepSeek V4 branch →
- PossiblePossibly related (embedding) · 50%Llama-Server is Throwing Away Your Perfectly Good KV Caches, and How to Fix It →
- PossiblePossibly related (embedding) · 61%FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference →
Covers
Implements
Related to
Covers (incoming)
newsI merged fixes for quantized KV cache into my DeepSeek V4 branchnewsLlama-Server is Throwing Away Your Perfectly Good KV Caches, and How to Fix ItnewsI developed my own quantized LLM from scratch, trained on 30B tokens, deploys in 60 MB [R]newsdeepseek-v4-flash-0731 - surprisingly usablenewsCachyLLama’s: llama.cpp fork with persistent KV cache that makes long local-agent sessions much less painfulnewsCachyLLama: llama.cpp fork with persistent SSD-backed KV caching for local agent workflowsnewsIs KV Cache in a high dimensional vector space? [D]newsDKV: Open-source KV-cache compression framework for local LLM inference (CLI + technical report)
Implements (incoming)
Related across the graph
paperA JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMsnewsCachyLLama’s: llama.cpp fork with persistent KV cache that makes long local-agent sessions much less painfulnewsWould having a dedicated programming language specifically for LLMs be a viable solution? [D]newsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)newsI developed my own quantized LLM from scratch, trained on 30B tokens, deploys in 60 MB [R]newsdeepseek-v4-flash-0731 - surprisingly usablepaperFreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM InferencenewsI merged fixes for quantized KV cache into my DeepSeek V4 branchnewsCachyLLama: llama.cpp fork with persistent SSD-backed KV caching for local agent workflowsnewsLlama-Server is Throwing Away Your Perfectly Good KV Caches, and How to Fix ItpaperRaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM InferencenewsDKV: Open-source KV-cache compression framework for local LLM inference (CLI + technical report)tutorialEvaluate a model properlynewsHardware startup unveils inference acceleratornewsIs KV Cache in a high dimensional vector space? [D]
