repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
aivrar/multi-turboquant
Unified KV cache compression for LLM inference — TurboQuant, IsoQuant, PlanarQuant, TriAttention. 10 methods, GPU-validated, multi-GPU planner. Compress KV cache 5-80x to run bigger models, longer context, more agents on your GPU.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 46%GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache →
- PossiblePossibly related (embedding) · 46%Going from single GPU to dual GPU is nice but not in the way I expected →
