Algebraic Decomposition Theory for Transformer Length Generalization
Transformer-based language models are known to sometimes generalize to sequences longer than seen during training, but we lack a precise characterization of which tasks admit length generalization. It is not even known which regular languages transformers length-generalize on -- and this is a foundational class of languages. Our contributions are to establish the first complete characterization of which regular languages transformers length-generalize on and provide a decision algorithm running in polynomial time in the size of the language's syntactic monoid. These results rely on an effectiv
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%Transformer →
- PossiblePossibly related (embedding) · 45%I built a compiler that turns computation graphs into the weights of a vanilla transformer — no training anywhere [P] →
- LinkedLinked via arxiv author · 85%Andy Yang →
“Algebraic Decomposition Theory for Transformer Length Generalization”
- LinkedLinked via arxiv author · 85%Blerta Veseli →
“Algebraic Decomposition Theory for Transformer Length Generalization”
- LinkedLinked via arxiv author · 85%Corentin Barloy →
“Algebraic Decomposition Theory for Transformer Length Generalization”
- LinkedLinked via arxiv author · 85%Michaël Cadilhac →
“Algebraic Decomposition Theory for Transformer Length Generalization”
- LinkedLinked via arxiv author · 85%Andreas Krebs →
“Algebraic Decomposition Theory for Transformer Length Generalization”
- LinkedLinked via arxiv author · 85%Charles Paperman →
“Algebraic Decomposition Theory for Transformer Length Generalization”
