Read original ↗
paperarXivTrust 82 · PrimaryPublished yesterdayLive · 1h ago

A Universal Context-Reuse Layer for Cross-Model KV Sharing

Modern large language model (LLM) serving systems increasingly operate over repeated or shared context, yet each model typically performs its own prefill computation even when another model has already processed the same input. Existing KV-cache reuse mechanisms substantially reduce redundant computation within a single model, but generally assume that the producer and consumer of a cache are identical. We study \emph{cross-model KV sharing}, which translates the KV state produced by a source model into a representation that can be consumed by a different target model, including models that di

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Yaoyi Li

    A Universal Context-Reuse Layer for Cross-Model KV Sharing

  • LinkedLinked via arxiv author · 85%Dongming Jiang

    A Universal Context-Reuse Layer for Cross-Model KV Sharing

  • LinkedLinked via arxiv author · 85%Tianyi Zhao

    A Universal Context-Reuse Layer for Cross-Model KV Sharing

  • LinkedLinked via arxiv author · 85%Bingzhe Li

    A Universal Context-Reuse Layer for Cross-Model KV Sharing

  • FuzzyOverlapping authors or contributors · 62%affaan-m/ECC

    Shared author/contributor keys: jiang

  • FuzzyOverlapping authors or contributors · 62%BerriAI/litellm

    Shared author/contributor keys: jiang

authored (incoming)

Implements (incoming)

Related across the graph

Topics