repoGitHubTrust 82 · PrimaryPublished 3mo agoLive · 1mo ago
attention-zoo
Implementations of many attention variants, benchmarked.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownSparse attention at million-token context →
- LinkedLinked via unknownBreakthrough in long-context efficiency announced →
- LinkedLinked via unknownAttention →
- LinkedLinked via unknownDnA: Denoising Attention for Visual Tasks →
- LinkedLinked via unknownNLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation →
- LinkedLinked via unknownPosition Bias Correction is Insufficient for One-Pass Attention Sorting →
- LinkedLinked via unknownMorphing into Hybrid Attention Models →
Implements
Covers
Related to
Implements (incoming)
paperDnA: Denoising Attention for Visual TaskspaperNLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window AdaptationpaperPosition Bias Correction is Insufficient for One-Pass Attention SortingpaperMorphing into Hybrid Attention ModelspaperEquiSteer: Cross-Attention Steering Towards a Fairer Text-Guided Image GenerationpaperUnderstanding Large Language ModelspaperSuper-Tuning: From Activation-Aware Pruning to Sparse Fine-TuningpaperInhibited Self-Attention: Sharpening Focus in Vision Transformers
Covers (incoming)
Related across the graph
paperNLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window AdaptationpaperPosition Bias Correction is Insufficient for One-Pass Attention SortingpaperInhibited Self-Attention: Sharpening Focus in Vision TransformerspaperSuper-Tuning: From Activation-Aware Pruning to Sparse Fine-TuningnewsHydraHead: From Head-Level Functional Heterogeneity to Specialized Attention Hybridization (from the Qwen team)paperMorphing into Hybrid Attention ModelspaperSparse attention at million-token contextpaperUnderstanding Large Language ModelsnewsBreakthrough in long-context efficiency announcedpaperDnA: Denoising Attention for Visual Tasksglossary_termAttentionpaperEquiSteer: Cross-Attention Steering Towards a Fairer Text-Guided Image Generation
