A Universal Context-Reuse Layer for Cross-Model KV Sharing
Modern large language model (LLM) serving systems increasingly operate over repeated or shared context, yet each model typically performs its own prefill computation even when another model has already processed the same input. Existing KV-cache reuse mechanisms substantially reduce redundant computation within a single model, but generally assume that the producer and consumer of a cache are identical. We study \emph{cross-model KV sharing}, which translates the KV state produced by a source model into a representation that can be consumed by a different target model, including models that di
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Yaoyi Li →
“A Universal Context-Reuse Layer for Cross-Model KV Sharing”
- LinkedLinked via arxiv author · 85%Dongming Jiang →
“A Universal Context-Reuse Layer for Cross-Model KV Sharing”
- LinkedLinked via arxiv author · 85%Tianyi Zhao →
“A Universal Context-Reuse Layer for Cross-Model KV Sharing”
- LinkedLinked via arxiv author · 85%Bingzhe Li →
“A Universal Context-Reuse Layer for Cross-Model KV Sharing”
- FuzzyOverlapping authors or contributors · 62%affaan-m/ECC →
“Shared author/contributor keys: jiang”
- FuzzyOverlapping authors or contributors · 62%BerriAI/litellm →
“Shared author/contributor keys: jiang”
