Read original ↗
paperarXivTrust 82 · PrimaryPublished 7d agoLive · 4d ago

BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference

Large Reasoning Models (LRMs) achieve superior problem-solving through extended Chain-of-Thought (CoT) generation, but the resulting key-value (KV) cache grows linearly with sequence length and creates severe memory bottlenecks, often exceeding GPU capacity for long reasoning traces. Existing KV cache compression methods rely on recent queries to estimate future token importance, implicitly assuming these serve as reliable proxies for future attention patterns. We demonstrate that this assumption fails in long-horizon reasoning: certain decoding steps generate Thought Revisiting Tokens (TRT) t

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 57%Language Models Can Control Their Own Attention [R]
  • FuzzySimilar title/name (fuzzy) · 87%LMCache/LMCache

    Fuzzy title match (0.94): “BeaconKV: Key-Value Cache Compression Guided by Beacon Queri” ≈ “LMCache/LMCache”

  • FuzzySimilar title/name (fuzzy) · 84%xorbitsai/inference

    Fuzzy title match (0.92): “BeaconKV: Key-Value Cache Compression Guided by Beacon Queri” ≈ “xorbitsai/inference”

  • FuzzyOverlapping authors or contributors · 62%ultralytics/yolov5

    Shared author/contributor keys: choi

  • LinkedLinked via arxiv author · 85%Janghyeon Kim

    BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference

  • LinkedLinked via arxiv author · 85%Minsoo Kim

    BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference

  • LinkedLinked via arxiv author · 85%Kyuhong Shim

    BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference

  • LinkedLinked via arxiv author · 85%Jungwook Choi

    BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference

Covers

Implements (incoming)

authored (incoming)

Related across the graph

Topics