Universal Transformers for Circuit Computations: Perfect Length Generalization in Tiny Transformers
Learning generalizable algorithmic computations remains a challenge for neural networks, as reflected in persistent failures on compositional and length generalization benchmarks. We present a provably correct, transformer parameterization (with only 280 learnable parameters for Boolean algebra tasks) capable of learning and evaluating problems of any depth or length. We assume inputs are fully parenthesized, well-formed expressions. Our approach conceptualizes algorithmic tasks as circuit models embedded in transformers, enabling depth-1 circuit reduction in a single forward pass. To achieve
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%I built a compiler that turns computation graphs into the weights of a vanilla transformer — no training anywhere [P] →
- PossiblePossibly related (embedding) · 49%I shrank a transformer until every number fitted on the screen and made the weights editable [R] →
- PossiblePossibly related (embedding) · 48%Transformer →
- PossiblePossibly related (embedding) · 47%Build your first transformer from scratch →
- PossiblePossibly related (embedding) · 49%A Mathematical Framework for Transformer Circuits (2021) →
- FuzzySimilar title/name (fuzzy) · 84%huggingface/transformers →
“Fuzzy title match (0.92): “Universal Transformers for Circuit Computations: Perfect Len” ≈ “huggingface/transformers””
- LinkedLinked via arxiv author · 85%Takuya Ito →
“Universal Transformers for Circuit Computations: Perfect Length Generalization in Tiny Transformers”
- LinkedLinked via arxiv author · 85%Ruchir Puri →
“Universal Transformers for Circuit Computations: Perfect Length Generalization in Tiny Transformers”
