newsReddit r/MachineLearningTrust 52 · CommunityPublished 6d agoLive · 4d ago
Language Models Can Control Their Own Attention [R]
Abstract Language models spend most of their attention on a small fraction of context, yet they read the entire KV cache to find the few tokens that matter. If the user asks about a previous detail in a 1M-token conversation, global attention layers must scan the full context to generate each token of the reply. A prominent approach mitigates this cost by pre-selecting relevant tokens via lightweight proxy scores, but this extrinsic scoring still in
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 92%Language Models Can Control Their Own Attention →
- PossiblePossibly related (embedding) · 73%Sparse attention at million-token context →
- PossiblePossibly related (embedding) · 70%Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads →
- PossiblePossibly related (embedding) · 67%ProxyFormer: A Dual-Stream Proxy Architecture for Ultra-Long Context and High-Resolution Generation →
- PossiblePossibly related (embedding) · 64%NLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation →
- PossiblePossibly related (embedding) · 57%BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference →
- PossiblePossibly related (embedding) · 45%Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time →
- PossiblePossibly related (embedding) · 55%Measuring LLM Sycophancy under Sustained Multi-Turn Pressure →
Covers
paperLanguage Models Can Control Their Own AttentionpaperSparse attention at million-token contextpaperLogit-Contribution Scoring Identifies Non-Literal Retrieval HeadspaperProxyFormer: A Dual-Stream Proxy Architecture for Ultra-Long Context and High-Resolution GenerationpaperNLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation
Covers (incoming)
paperBeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model InferencepaperInfluence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference timepaperMeasuring LLM Sycophancy under Sustained Multi-Turn PressurepaperIt's Not RoPE that Creates Sinks: The Role of Self-Concentration and Value-Non-Mixing in Attention
Related across the graph
paperNLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window AdaptationpaperLanguage Models Can Control Their Own AttentionpaperLogit-Contribution Scoring Identifies Non-Literal Retrieval HeadspaperInfluence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference timepaperSparse attention at million-token contextpaperProxyFormer: A Dual-Stream Proxy Architecture for Ultra-Long Context and High-Resolution GenerationpaperBeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model InferencepaperIt's Not RoPE that Creates Sinks: The Role of Self-Concentration and Value-Non-Mixing in AttentionpaperMeasuring LLM Sycophancy under Sustained Multi-Turn Pressure
