Mobius Learning: Cyclic Depth Folding in Transformers
Transformer-based language models organize computation along an ordered depth axis, where shallow and deep blocks often develop distinct representational roles. We challenge the conventional view that these roles must remain tied to a block's position in the ordered sequence. We introduce Mobius Learning, a training architecture based on cyclic depth folding, in which different data streams follow cyclically shifted block orders. The same block group is therefore applied early in the block sequence for some data streams and late for others, so it is optimized in both shallow and deep roles, a
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 56%Transformer →
- PossiblePossibly related (embedding) · 45%Transformers in Deep Learning: How Self-Attention Changed Modern AI - Snowflake →
- FuzzySimilar title/name (fuzzy) · 87%lucidrains/x-transformers →
“Fuzzy title match (0.94): “Mobius Learning: Cyclic Depth Folding in Transformers” ≈ “lucidrains/x-transformers””
- FuzzySimilar title/name (fuzzy) · 84%huggingface/transformers →
“Fuzzy title match (0.92): “Mobius Learning: Cyclic Depth Folding in Transformers” ≈ “huggingface/transformers””
- FuzzySimilar title/name (fuzzy) · 59%aymericdamien/TopDeepLearning →
“Fuzzy title match (0.73): “Mobius Learning: Cyclic Depth Folding in Transformers” ≈ “aymericdamien/TopDeepLearning””
- LinkedLinked via arxiv author · 85%Tongtian Zhu →
“Mobius Learning: Cyclic Depth Folding in Transformers”
