newsReddit r/MachineLearningTrust 52 · CommunityPublished yesterdayLive · 9h ago
Sliding-window attention beats linear on long-context reasoning [R]
Sliding Window Attention with sinks, one of the simplest existing fixes for the quadratic-cost problem in LLMs, holds up as well or better than the linear-attention variants labs have been spending post-training compute to produce. That is the claim of a [new arXiv preprint]( https://arxiv.org/abs/2608.28444 ) by Alexia Jolicoeur-Martineau, Rhea Sanjay Sukthanker, Pashmina Cameron and Emy Gervais. On the long-c
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 65%Sliding-window beats linear attention →
- PossiblePossibly related (embedding) · 52%NLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation →
- PossiblePossibly related (embedding) · 49%Semantic Head Specialization Guides Hybrid ViT Attention for Multimodal LLMs →
- PossiblePossibly related (embedding) · 48%Sparse attention at million-token context →
- PossiblePossibly related (embedding) · 48%lucidrains/taylor-series-linear-attention →
Covers
paperSliding-window beats linear attentionpaperNLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window AdaptationpaperSemantic Head Specialization Guides Hybrid ViT Attention for Multimodal LLMspaperSparse attention at million-token contextrepolucidrains/taylor-series-linear-attention
Related across the graph
paperNLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptationrepolucidrains/taylor-series-linear-attentionpaperSparse attention at million-token contextpaperSemantic Head Specialization Guides Hybrid ViT Attention for Multimodal LLMspaperSliding-window beats linear attention
