DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression
Long-context language model inference is increasingly limited by the memory bandwidth and capacity required to store key-value caches, yet existing compression methods often apply uniform budgets across layers or tokens and degrade retrieval when lexical cues and semantic states require different preservation. We introduce DepthWeave-KV, a token-adaptive cache compression method that factorizes key and value states across neighboring transformer layers using shared low-rank channel bases while retaining lightweight token-specific residuals where attention behavior is sensitive. DepthWeave-KV c
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%Zefan-Cai/R-KV →
- PossiblePossibly related (embedding) · 52%Breakthrough in long-context efficiency announced →
- PossiblePossibly related (embedding) · 48%Zefan-Cai/KVCache-Factory →
- PossiblePossibly related (embedding) · 47%Transformer →
- LinkedLinked via arxiv author · 85%Anna Cordoba →
“DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression”
- LinkedLinked via arxiv author · 85%Adam Puente Tercero →
“DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression”
- LinkedLinked via arxiv author · 85%Nerea Angulo Hijo →
“DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression”
- LinkedLinked via arxiv author · 85%Mar Linares Tercero →
“DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression”
