Read original ↗
newsAI NewsTrust 60Published 1mo agoLive · 2mo ago

Breakthrough in long-context efficiency announced

A new attention scheme cuts memory use for very long inputs.

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Covers (incoming)

paperCARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attentionrepoattention-zoopaperNLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window AdaptationpaperPosition Bias Correction is Insufficient for One-Pass Attention SortingpaperVisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual ContextpaperLLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via a Proprioceptive DashboardpaperMorphing into Hybrid Attention ModelspaperERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMspaperRaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM InferencepaperLogit-Contribution Scoring Identifies Non-Literal Retrieval HeadspaperUnderstanding Large Language ModelspaperA Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State ForgetspaperMemDefrag: Latent Memory Defragmentation for Large Language ModelspaperFreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM InferencepaperDepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache CompressionpaperSuper-Tuning: From Activation-Aware Pruning to Sparse Fine-TuningpaperSelf-Guided Test-Time Training for Long-Context LLMspaperExtending LLM Context via Associative Recurrent MemorypaperLong-Context Fine-Tuning with Limited VRAMpaperGradient Concentration, Not Weight Saliency, Explains Representation-Level Class Unlearning

Related across the graph

paperNLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window AdaptationpaperPosition Bias Correction is Insufficient for One-Pass Attention SortingpaperSuper-Tuning: From Activation-Aware Pruning to Sparse Fine-TuningpaperGradient Concentration, Not Weight Saliency, Explains Representation-Level Class UnlearningpaperDepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache CompressionpaperLong-Context Fine-Tuning with Limited VRAMpaperExtending LLM Context via Associative Recurrent MemorypaperLogit-Contribution Scoring Identifies Non-Literal Retrieval HeadspaperERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMspaperLLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via a Proprioceptive DashboardpaperMorphing into Hybrid Attention ModelspaperCARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear AttentionpaperFreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM InferencepaperSparse attention at million-token contextpaperA Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State ForgetspaperVisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual ContextpaperUnderstanding Large Language ModelspaperRaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM InferencepaperMemDefrag: Latent Memory Defragmentation for Large Language Modelsglossary_termAttentionpaperSelf-Guided Test-Time Training for Long-Context LLMsrepoattention-zoo