FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation
Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, but adaptive Top-$p$ routing creates uneven per-head workloads under multi-GPU sequence parallelism. The resulting workload heterogeneity turns sparse attention into a rank-level straggler problem. We present \method{}, a training-free sparse-attention system that improves the distributed execution efficiency of adaptive sparse attention under multi-GPU sequence parallelism. \method{} uses Top-$p$ r
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 47%Reducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading - NVIDIA Developer →
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzySimilar title/name (fuzzy) · 59%Developer-Y/cs-video-courses →
“Fuzzy title match (0.73): “FVAttn: Adaptive Sparse Attention with Runtime Load Balancin” ≈ “Developer-Y/cs-video-courses””
- LinkedLinked via arxiv author · 85%Jihao Liu →
“FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation”
- LinkedLinked via arxiv author · 85%Chenghuan Huang →
“FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation”
- LinkedLinked via arxiv author · 85%Ye Huang →
“FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation”
- LinkedLinked via arxiv author · 85%Zhiying Wen →
“FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation”
- LinkedLinked via arxiv author · 85%Mohan Zhang →
“FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation”
