HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation
Autoregressive (AR) video diffusion models have become a promising paradigm for long and streaming video synthesis, but the continuously growing Key-Value (KV) cache makes attention the dominant inference cost, especially at high resolution where each frame contributes many tokens. Existing remedies either evict the cache with coarse heuristics that cause inter-frame flickering, or require model re-training. We propose HeadCast, a training-free, plug-and-play acceleration framework built on the observation that a pre-trained AR model's attention heads exhibit stable, heterogeneous behaviors. A
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%affaan-m/ECC →
“Shared author/contributor keys: jiang”
- FuzzyOverlapping authors or contributors · 62%BerriAI/litellm →
“Shared author/contributor keys: jiang”
- FuzzyOverlapping authors or contributors · 62%deepspeedai/DeepSpeed →
“Shared author/contributor keys: lai”
- FuzzySimilar title/name (fuzzy) · 59%Developer-Y/cs-video-courses →
“Fuzzy title match (0.73): “HeadCast: Casting Attention Heads for Efficient Autoregressi” ≈ “Developer-Y/cs-video-courses””
- LinkedLinked via arxiv author · 85%Jinliang Shen →
“HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation”
- LinkedLinked via arxiv author · 85%Lianghao Su →
“HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation”
- LinkedLinked via arxiv author · 85%Zheming Li →
“HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation”
- LinkedLinked via arxiv author · 85%Kang He →
“HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation”
