Skip to main content
Angestrom home
SearchPapersModelsLive AIIntelligence
Search⌕⌘K
EnterprisePricingSign in

Stay Ahead in the AI Revolution

Weekly digest — EPI pulse, top intelligence, fresh lineage. Free, no account.

Follow Angestrom
Global source network
Synced every 5 minutes

Continuous sync from primary AI sources — indexed, enriched, and queryable in real time.

arXivHugging FaceGitHubOpenAIAnthropicDeepMindReutersBBC TechHacker NewsReddit MLVerified feedsFunding
ANGESTROM

The Intelligence Layer of Humanity. Everything AI. All in One Place.

Angestrom connects every piece of the AI ecosystem — data, models, research, companies, tools, and people.

info@angestrom.comwww.angestrom.comLucknow, Uttar Pradesh, India

Product

  • AI Search
  • AI Models
  • Research Papers
  • Companies
  • News & Events
  • GitHub Explorer
  • APIs & Tools
  • Datasets
  • Benchmarks
  • Model lifecycle
  • Funding graph
  • Contributors
  • AI Agents

Resources

  • Weekly digest
  • Documentation
  • Tutorials
  • Guides
  • News
  • Help / Start
  • Community

Company

  • About
  • Contact
  • Privacy Policy
  • Terms of Service
  • Acceptable Use

Enterprise

  • Pricing
  • Workspace
  • Contact Sales

Developer

  • Developer Hub
  • API docs
  • GitHub

Learn

  • Learning Academy
  • Roadmaps
  • Glossary
  • AI for Beginners

Popular Topics

Loading topics…
View All Topics →
© 2026 Angestrom Intelligence Private Limited. All rights reserved.
English
Theme
Angestrom home
SearchPapersModelsLive AIIntelligence
Search⌕⌘K
EnterprisePricingSign in
  1. Home
  2. /Repositories
  3. /LMCache/LMCache
Read original ↗
repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · yesterday

LMCache/LMCache

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 50%Biggest, baddest model to fill 144GB VRAM + 120GB RAM to the brim, regardless of speed →
  • PossiblePossibly related (embedding) · 58%I mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset) →
  • PossiblePossibly related (embedding) · 49%I compiled LLM inference pricing across 7 providers — the caching numbers are surprising(spreadsheet included) [R] →
  • PossiblePossibly related (embedding) · 48%One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining →
  • PossiblePossibly related (embedding) · 47%Devs - you have 64gb of VRAM - which model do you use for coding? →
  • FuzzySimilar title/name (fuzzy) · 87%A JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMs →

    “Fuzzy title match (0.94): “A JoLT for the KV Cache: Near-Lossless KV Cache Compression ” ≈ “LMCache/LMCache””

  • FuzzySimilar title/name (fuzzy) · 87%DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression →

    “Fuzzy title match (0.94): “DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factoriza” ≈ “LMCache/LMCache””

  • FuzzySimilar title/name (fuzzy) · 87%Evolving Cache Schedules for Fast Diffusion Policy Inference →

    “Fuzzy title match (0.94): “Evolving Cache Schedules for Fast Diffusion Policy Inference” ≈ “LMCache/LMCache””

Covers

newsBiggest, baddest model to fill 144GB VRAM + 120GB RAM to the brim, regardless of speednewsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)newsI compiled LLM inference pricing across 7 providers — the caching numbers are surprising(spreadsheet included) [R]newsDevs - you have 64gb of VRAM - which model do you use for coding?

Implements

paperOne-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM PretrainingpaperA JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMspaperDepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache CompressionpaperEvolving Cache Schedules for Fast Diffusion Policy InferencepaperError Certificates for KV-Cache Eviction via Randomized Design

Covers (incoming)

newsI merged fixes for quantized KV cache into my DeepSeek V4 branchnewsLlama-Server is Throwing Away Your Perfectly Good KV Caches, and How to Fix ItnewsReducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading - NVIDIA Developer

Related to (incoming)

paperFreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM InferencepaperPagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization

Related across the graph

newsReducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading - NVIDIA DeveloperpaperEvolving Cache Schedules for Fast Diffusion Policy InferencenewsI compiled LLM inference pricing across 7 providers — the caching numbers are surprising(spreadsheet included) [R]paperA JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMspaperDepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache CompressionpaperError Certificates for KV-Cache Eviction via Randomized DesignnewsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)paperFreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM InferencenewsDevs - you have 64gb of VRAM - which model do you use for coding?newsI merged fixes for quantized KV cache into my DeepSeek V4 branchnewsLlama-Server is Throwing Away Your Perfectly Good KV Caches, and How to Fix ItnewsBiggest, baddest model to fill 144GB VRAM + 120GB RAM to the brim, regardless of speedpaperPagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight QuantizationpaperOne-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining
Knowledge path·NReducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading - NVIDIA Developer→PEvolving Cache Schedules for Fast Diffusion Policy Inference→NI compiled LLM inference pricing across 7 providers — the caching numbers are surprising(spreadsheet included) [R]→RLMCache/LMCache

Topics

amdcudafastinferencekv-cachellmpytorchrocmspeedvllm

Explore

Search similar →Knowledge graph →All repos →Full intelligence feed →
Graph trust82Primary
Graph score11166