Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

Self-Gating Attention for Efficient Time Series Forecasting

Transformer architectures have shown strong potential in time series forecasting, where multi-head self-attention is widely used to capture temporal dependencies across historical timestamps. However, standard self-attention has quadratic time and memory complexity with respect to the look-back length. This cost may limit its use in resource-constrained or high-throughput forecasting systems, where fast and memory-efficient inference is important. Through qualitative and quantitative analyses, we observe that self-attention maps in time series forecasting often contain redundant patterns acros

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 54%amazon-science/chronos-forecasting
  • PossiblePossibly related (embedding) · 47%Nixtla/statsforecast
  • LinkedLinked via arxiv author · 85%Dezheng Wang

    Self-Gating Attention for Efficient Time Series Forecasting

  • LinkedLinked via arxiv author · 85%Tong Chen

    Self-Gating Attention for Efficient Time Series Forecasting

  • LinkedLinked via arxiv author · 85%Wei Yuan

    Self-Gating Attention for Efficient Time Series Forecasting

  • LinkedLinked via arxiv author · 85%Congyan Chen

    Self-Gating Attention for Efficient Time Series Forecasting

  • LinkedLinked via arxiv author · 85%Shihua Li

    Self-Gating Attention for Efficient Time Series Forecasting

  • LinkedLinked via arxiv author · 85%Hongzhi Yin

    Self-Gating Attention for Efficient Time Series Forecasting

Implements

authored (incoming)

Implements (incoming)

Covers (incoming)

Related across the graph

Topics