repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 26d ago
novitalabs/pegaflow
High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%I mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset) →
- PossiblePossibly related (embedding) · 53%OpenAI and Broadcom announce chip designed for LLM inference at scale →
- PossiblePossibly related (embedding) · 52%Tesla V100 16GB local LLMs, single and dual NVLink benchmarks →
- PossiblePossibly related (embedding) · 50%WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs →
- PossiblePossibly related (embedding) · 48%OpenAI and Broadcom unveil LLM-optimized inference chip →
- PossiblePossibly related (embedding) · 49%Llama-Server is Throwing Away Your Perfectly Good KV Caches, and How to Fix It →
- PossiblePossibly related (embedding) · 50%Google Cloud's Always-On Memory Agent Replaces RAG and Embeddings With Continuous LLM Consolidation on Gemini 3.1 Flash-Lite - MarkTechPost →
- PossiblePossibly related (embedding) · 54%CachyLLama: llama.cpp fork with persistent SSD-backed KV caching for local agent workflows →
Covers
Implements
Covers (incoming)
newsLlama-Server is Throwing Away Your Perfectly Good KV Caches, and How to Fix ItnewsGoogle Cloud's Always-On Memory Agent Replaces RAG and Embeddings With Continuous LLM Consolidation on Gemini 3.1 Flash-Lite - MarkTechPostnewsCachyLLama: llama.cpp fork with persistent SSD-backed KV caching for local agent workflows
Related across the graph
newsOpenAI and Broadcom announce chip designed for LLM inference at scalenewsOpenAI and Broadcom unveil LLM-optimized inference chipnewsGoogle Cloud's Always-On Memory Agent Replaces RAG and Embeddings With Continuous LLM Consolidation on Gemini 3.1 Flash-Lite - MarkTechPostnewsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)paperWattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMsnewsTesla V100 16GB local LLMs, single and dual NVLink benchmarksnewsCachyLLama: llama.cpp fork with persistent SSD-backed KV caching for local agent workflowsnewsLlama-Server is Throwing Away Your Perfectly Good KV Caches, and How to Fix It
