Morphing into Hybrid Attention Models
Hybrid attention models improve long-context efficiency by retaining only a subset of full-attention layers and replacing the remaining layers with linear attention. However, the effectiveness of Transformer-to-hybrid conversion critically depends on which layers preserve full attention. Existing hybrid layer selection methods typically rely on heuristic strategies such as fixed placement patterns or layerwise scoring, implicitly treating layer importance as isolated and overlooking the interdependent layer effect under a global hybrid configuration. In this work, we formulate hybrid layer sel
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownattention-zoo →
- LinkedLinked via unknownAttention →
- LinkedLinked via unknownBreakthrough in long-context efficiency announced →
- LinkedLinked via unknownlucidrains/poly-attention →
- LinkedLinked via unknownTransformer →
- PossiblePossibly related (embedding) · 47%lucidrains/x-transformers →
- PossiblePossibly related (embedding) · 55%Introducing DWARF-55M-Base →
