Cac
44 items across the graph · 2 news stories — tagged with Cac.
Latest news
The End of the Single-Model AI Era - Communications of the ACM
The End of the Single-Model AI Era Communications of the ACM
Read full story →More news · 1
From the graph · 42
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
14MB foundation model for tiny devices; phones, wearables, smart home, and robots.
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
Open-source multimodal retrieval engine (Morphik Core). By Morphik — AI back office for skilled nursing & senior living (morphik.ai).
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
Unified KV Cache Compression Methods for Auto-Regressive Models
[Neurips 2025] R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
LLM KV cache compression made easy
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!
Shared Single-file memory layer for all your agents, sub mili-second RAG over text, photo and video on Apple Silicon.. No Server. No API. One File. Pure Swift
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Intelligent token optimization for Claude Code - achieving 95%+ token reduction through caching, compression, and smart tool intelligence
Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default, cloud by choice. Offline AI memory (taOSmd), self-ho…
Harness-oriented agent system framework for production-grade LLM agent applications
Fixes prompt cache regression in Claude Code that causes up to 20x cost increase on resumed sessions
Redis Vector Library (RedisVL) -- the AI-native Python client for Redis.
Persist and reuse KV Cache to speedup your LLM.
CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system effici…
Suite of tools containing an in-memory vector datastore and AI proxy
Alibaba Cloud's high-performance KVCache system for LLM inference, with components for global cache management, inference simulation(HiSim), and more.
High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.
Anthropic Claude API wrapper for Go
CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system effici…
Production LLM call layer for AI agents and tools: keep OpenAI/Anthropic/AI SDK/LiteLLM, hot-swap models with MDA presets, and add cache, retries, circuit break…
elearning linux/mac/db/cache/server/tools/人工智能/安全防护/黑科技
A lean, fully-tested LLM inference server for the hardware you already have — free tiers, shared VPS, 2-core ARM boxes. OpenAI-compatible API on llama.cpp. On a…
Extreme weight + KV cache compression for LLMs on Apple Silicon (MLX implementation of Google's TurboQuant)
Reliable and Efficient Semantic Prompt Caching with vCache
为 DeepSeek 前缀缓存定制的终端 Code Agent(纯 Go),缓存命中率 95-99%,输入成本降至 1/50。A terminal coding agent optimized for DeepSeek prefix caching — 95-99% cache hit, 1/50th the cost…
High-Performance KV Cache Storage Engine on CXL Shared Memory for LLM Inference
linked of Romanian fiscal authority(MFP-ANAF-GOV-RO)
ConDB: The KV-Cache Native Context Database
RAMen is a fast in-memory data store like Redis, but built for AI: drop-in Redis protocol, native vector search, semantic caching, and a built-in MCP server for…
Native Windows builds of vLLM 0.25.1 with CUDA 12.8, Python 3.13, Multi-TurboQuant, and experimental CPU/NVMe KV-cache offload.
qMLX: Custom inference engine for Qwen 3.5 122B on Apple Silicon, extending MLX with hybrid attention support, SSD-backed KV cache, and RYS layer duplication fo…
A curated list of strategies, tools, papers, and resources for reducing LLM token costs and improving efficiency in production.
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures valid…
Make local LLM inference faster with chunk-level KV cache reuse
Unified KV cache compression for LLM inference — TurboQuant, IsoQuant, PlanarQuant, TriAttention. 10 methods, GPU-validated, multi-GPU planner. Compress KV cach…
Minimal, zero-dependency LLM inference in pure C11. CPU-first with NEON/AVX2 SIMD. Flash MoE (pread + LRU expert cache). TurboQuant 3-bit KV compression (8.9x l…
