Semantic Head Specialization Guides Hybrid ViT Attention for Multimodal LLMs
Hybrid attention dominates frontier LLMs, yet Vision Transformers (ViTs) in multimodal LLMs lack a satisfactory hybrid design, with no consensus on why certain attention patterns work better. To fill this gap, we study ViT attention heads and find they differentiate into object- and background-specialist roles, a pattern most pronounced under full attention; we call this Semantic Head Specialization (SHS). We propose SHS-Index to quantify this specialization, show that it distinguishes full-attention from chunk-window ViTs, and find that it strongly tracks downstream benchmark performance. We
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 57%Breakthrough in long-context efficiency announced →
- PossiblePossibly related (embedding) · 53%HydraHead: From Head-Level Functional Heterogeneity to Specialized Attention Hybridization (from the Qwen team) →
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%Kong/kong →
“Shared author/contributor keys: kong”
- FuzzySimilar title/name (fuzzy) · 59%microsoft/semantic-kernel →
“Fuzzy title match (0.73): “Semantic Head Specialization Guides Hybrid ViT Attention for” ≈ “microsoft/semantic-kernel””
- LinkedLinked via arxiv author · 85%Chenhong He →
“Semantic Head Specialization Guides Hybrid ViT Attention for Multimodal LLMs”
- LinkedLinked via arxiv author · 85%Lei Li →
“Semantic Head Specialization Guides Hybrid ViT Attention for Multimodal LLMs”
- LinkedLinked via arxiv author · 85%Shicheng Li →
“Semantic Head Specialization Guides Hybrid ViT Attention for Multimodal LLMs”
