repoGitLabTrust 82 · PrimaryPublished 28d agoLive · 25d ago
pubmethod/quantization
Following Google's TurboQuant release (https://research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/), this repository contains two KV cache quantization simulations (KMeans codebook and TurboQuant-style rotation) to explore core quantization principles. **Inspired by**: https://medium.com/@gautsoni/llm-quantization-the-practical-guide-and-why-it-matters-for-inference-and-training-8668f4b91dcc
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%QuantBench →
