Skip to main content
Angestrom home
SearchPapersModelsLive AIIntelligence
Search⌕⌘K
EnterprisePricingSign in

Stay Ahead in the AI Revolution

Weekly digest — EPI pulse, top intelligence, fresh lineage. Free, no account.

Follow Angestrom
Global source network
Synced every 5 minutes

Continuous sync from primary AI sources — indexed, enriched, and queryable in real time.

arXivHugging FaceGitHubOpenAIAnthropicDeepMindReutersBBC TechHacker NewsReddit MLVerified feedsFunding
Angestrom

Angestrom connects every piece of the AI ecosystem — data, models, research, companies, tools, and people.

info@angestrom.comwww.angestrom.comLucknow, Uttar Pradesh, India

Product

  • AI Search
  • AI Models
  • Research Papers
  • Companies
  • News & Events
  • GitHub Explorer
  • APIs & Tools
  • Datasets
  • Benchmarks
  • Model lifecycle
  • Funding graph
  • Contributors
  • AI Agents

Resources

  • Weekly digest
  • Documentation
  • Tutorials
  • Guides
  • News
  • Help / Start
  • Community

Company

  • About
  • Contact
  • Privacy Policy
  • Terms of Service
  • Acceptable Use

Enterprise

  • Pricing
  • Workspace
  • Contact Sales

Developer

  • Developer Hub
  • API docs
  • GitHub

Learn

  • Learning Academy
  • Roadmaps
  • Glossary
  • AI for Beginners

Popular Topics

Loading topics…
View All Topics →
© 2026 Angestrom. All rights reserved.
English
Theme
Angestrom home
SearchPapersModelsLive AIIntelligence
Search⌕⌘K
EnterprisePricingSign in
  1. Home
  2. /Repositories
  3. /pythongiant/KVBoost
Read original ↗
repoGitHubTrust 82 · PrimaryPublished 2mo agoLive · 2mo ago

pythongiant/KVBoost

Make local LLM inference faster with chunk-level KV cache reuse

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 56%I mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset) →
  • PossiblePossibly related (embedding) · 51%Hardware startup unveils inference accelerator →
  • PossiblePossibly related (embedding) · 49%Would having a dedicated programming language specifically for LLMs be a viable solution? [D] →
  • PossiblePossibly related (embedding) · 48%RaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM Inference →
  • PossiblePossibly related (embedding) · 48%Evaluate a model properly →
  • PossiblePossibly related (embedding) · 45%I merged fixes for quantized KV cache into my DeepSeek V4 branch →
  • PossiblePossibly related (embedding) · 50%Llama-Server is Throwing Away Your Perfectly Good KV Caches, and How to Fix It →
  • PossiblePossibly related (embedding) · 61%FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference →

Covers

newsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)newsHardware startup unveils inference acceleratornewsWould having a dedicated programming language specifically for LLMs be a viable solution? [D]

Implements

paperRaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM Inference

Related to

tutorialEvaluate a model properly

Covers (incoming)

newsI merged fixes for quantized KV cache into my DeepSeek V4 branchnewsLlama-Server is Throwing Away Your Perfectly Good KV Caches, and How to Fix ItnewsI developed my own quantized LLM from scratch, trained on 30B tokens, deploys in 60 MB [R]newsdeepseek-v4-flash-0731 - surprisingly usablenewsCachyLLama’s: llama.cpp fork with persistent KV cache that makes long local-agent sessions much less painfulnewsCachyLLama: llama.cpp fork with persistent SSD-backed KV caching for local agent workflowsnewsIs KV Cache in a high dimensional vector space? [D]newsDKV: Open-source KV-cache compression framework for local LLM inference (CLI + technical report)

Implements (incoming)

paperFreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM InferencepaperA JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMs

Related across the graph

paperA JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMsnewsCachyLLama’s: llama.cpp fork with persistent KV cache that makes long local-agent sessions much less painfulnewsWould having a dedicated programming language specifically for LLMs be a viable solution? [D]newsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)newsI developed my own quantized LLM from scratch, trained on 30B tokens, deploys in 60 MB [R]newsdeepseek-v4-flash-0731 - surprisingly usablepaperFreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM InferencenewsI merged fixes for quantized KV cache into my DeepSeek V4 branchnewsCachyLLama: llama.cpp fork with persistent SSD-backed KV caching for local agent workflowsnewsLlama-Server is Throwing Away Your Perfectly Good KV Caches, and How to Fix ItpaperRaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM InferencenewsDKV: Open-source KV-cache compression framework for local LLM inference (CLI + technical report)tutorialEvaluate a model properlynewsHardware startup unveils inference acceleratornewsIs KV Cache in a high dimensional vector space? [D]
Knowledge path·PA JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMs→NCachyLLama’s: llama.cpp fork with persistent KV cache that makes long local-agent sessions much less painful→NWould having a dedicated programming language specifically for LLMs be a viable solution? [D]→Rpythongiant/KVBoost

Topics

kv-cachekv-cache-lpllmllm-inferencellm-optimizationlocal-ailocal-ai-llmlocal-llmopen-llm

Explore

Search similar →Knowledge graph →All repos →Full intelligence feed →
Graph trust82Primary
Graph score25