repoGitHubTrust 82 · PrimaryPublished 2d agoLive · 2d ago
NVIDIA/kvpress
LLM KV cache compression made easy
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 67%DKV: Open-source KV-cache compression framework for local LLM inference (CLI + technical report) →
- PossiblePossibly related (embedding) · 53%Reducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading - NVIDIA Developer →
- PossiblePossibly related (embedding) · 52%I mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset) →
