SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while retaining the favorable long-sequence scaling of linear attention. To avoid quadratic attention throughout, Hybrid Linear-Softmax Attention combines gated linear attention for O(N)-dominated mixing with periodic gated-softmax anchors at a 3:1 ratio, restoring the full-rank token interactions that pure linear attention lacks. To propagat
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 78%sgl-project/sglang →
“Shared author/contributor keys: luo, zhou”
- FuzzyOverlapping authors or contributors · 62%janhq/jan →
“Shared author/contributor keys: han”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%ultralytics/ultralytics →
“Shared author/contributor keys: han”
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- LinkedLinked via arxiv author · 85%Junsong Chen →
“SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation”
- LinkedLinked via arxiv author · 85%Jincheng Yu →
“SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation”
- LinkedLinked via arxiv author · 85%Yitong Li →
“SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation”
