repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 29d ago
Zefan-Cai/R-KV
[Neurips 2025] R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 47%New benchmark exposes reasoning gaps in top models →
- PossiblePossibly related (embedding) · 46%Retrace-1.5B →
- PossiblePossibly related (embedding) · 46%DeepSeek-V4-Flash (MXFP4): compute buffer scales ~3x just from KV cache quant type (f16 vs q8_0) — anyone else seeing this? Llama.cpp →
- PossiblePossibly related (embedding) · 45%CARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attention →
- PossiblePossibly related (embedding) · 45%Message Passing Enables Efficient Reasoning →
- PossiblePossibly related (embedding) · 49%CheckRLM: Effective Knowledge-Thought Coherence Checking in Retrieval-Augmented Reasoning →
- PossiblePossibly related (embedding) · 54%DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression →
- PossiblePossibly related (embedding) · 51%Byte exact KV cache grafting on frozen Gemma 4 →
Covers
Related to
Implements
Implements (incoming)
Covers (incoming)
Related across the graph
paperCheckRLM: Effective Knowledge-Thought Coherence Checking in Retrieval-Augmented ReasoningpaperDepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache CompressionnewsDeepSeek-V4-Flash (MXFP4): compute buffer scales ~3x just from KV cache quant type (f16 vs q8_0) — anyone else seeing this? Llama.cpppaperMessage Passing Enables Efficient ReasoningpaperCARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear AttentionmodelRetrace-1.5BnewsByte exact KV cache grafting on frozen Gemma 4newsNew benchmark exposes reasoning gaps in top models
