newsReddit r/MachineLearningTrust 52 · CommunityPublished 2d agoLive · 5h ago
How can we solve long-range recall in linear attention? [D]
Recently, I started working on DNA sequence modeling and decided to explore linear attention , mainly because DNA sequences can easily reach 1M tokens , making standard softmax attention extremely expensive in terms of memory and computation. The model performed reasonably well on several benchmarks, but I ran into a major problem with long-range recall . On a Needle in a Haystack-style benchmark, my model w
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 59%DnA: Denoising Attention for Visual Tasks →
- PossiblePossibly related (embedding) · 55%A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets →
- PossiblePossibly related (embedding) · 51%RaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM Inference →
- PossiblePossibly related (embedding) · 51%Sparse attention at million-token context →
- PossiblePossibly related (embedding) · 50%Long-Context Fine-Tuning with Limited VRAM →
Covers
paperDnA: Denoising Attention for Visual TaskspaperA Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State ForgetspaperRaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM InferencepaperSparse attention at million-token contextpaperLong-Context Fine-Tuning with Limited VRAM
Related across the graph
paperLong-Context Fine-Tuning with Limited VRAMpaperSparse attention at million-token contextpaperA Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State ForgetspaperRaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM InferencepaperDnA: Denoising Attention for Visual Tasks
