It's Not RoPE that Creates Sinks: The Role of Self-Concentration and Value-Non-Mixing in Attention
Large Language Models (LLMs) often exhibit "Attention Sink" (AS) and the accompanying "Massive Activations" (MAs) at the initial position of a sequence. These phenomena frequently co-occur, and MAs can pose challenges for low-bit quantization. In this study, we analyze the factors underlying AS and MAs that emerge at the initial position regardless of the token occupying it. Our experiments suggest that self-concentration of attention, resulting from the causal mask, and the subsequent Value-non-mixing in attention outputs contribute to AS and MAs. These findings provide new empirical evidence
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 63%Language Models Can Control Their Own Attention [R] →
- PossiblePossibly related (embedding) · 55%Attention →
- PossiblePossibly related (embedding) · 51%Understanding large language models demands distinguishing human projection from machine cognition - Nature →
- PossiblePossibly related (embedding) · 50%Breakthrough in long-context efficiency announced →
- PossiblePossibly related (embedding) · 49%Study: Generative AI is making writing on Reddit and elsewhere boring →
- LinkedLinked via arxiv author · 85%Raito Kiya →
“It's Not RoPE that Creates Sinks: The Role of Self-Concentration and Value-Non-Mixing in Attention”
- LinkedLinked via arxiv author · 85%Satoki Ohashi →
“It's Not RoPE that Creates Sinks: The Role of Self-Concentration and Value-Non-Mixing in Attention”
- LinkedLinked via arxiv author · 85%Kosuke Sato →
“It's Not RoPE that Creates Sinks: The Role of Self-Concentration and Value-Non-Mixing in Attention”
