The State-Prediction Separation Hypothesis
Transformers use the same forward computation stream to both predict the next token and store useful state for future token predictions. We formulate the \emph{state-prediction separation hypothesis}: disentangling the two roles yields better language modeling performance. We design a Transformer variant that uses two computation streams to separate the two functions, and conduct pretraining experiments across various scales. Our experiments show that state-prediction separation consistently offers better data and compute efficiencies, improving validation loss and outperforming standard Trans
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownTransformer →
- LinkedLinked via unknownengineering87/llm-atlas →
- LinkedLinked via unknownBuild your first transformer from scratch →
- LinkedLinked via unknownquant-kit →
- LinkedLinked via arxiv author · 85%Giovanni Monea →
“The State-Prediction Separation Hypothesis”
- LinkedLinked via arxiv author · 85%Nathan Godey →
“The State-Prediction Separation Hypothesis”
- LinkedLinked via arxiv author · 85%Kianté Brantley →
“The State-Prediction Separation Hypothesis”
- LinkedLinked via arxiv author · 85%Yoav Artzi →
“The State-Prediction Separation Hypothesis”
