tutorialAngestrom AcademyTrust 60Published 3mo agoLive · 3mo ago
Build your first transformer from scratch
Step through attention, MLPs, and training on a toy task.
Step through attention, MLPs, and training on a toy task. Step through attention, MLPs, and training on a toy task.
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownGrokking in small transformers →
- LinkedLinked via unknownquant-kit →
- LinkedLinked via unknownGeneralization Analysis of Transformers in Distribution Regression →
- LinkedLinked via unknownThe State-Prediction Separation Hypothesis →
- PossiblePossibly related (embedding) · 58%lucidrains/x-transformers →
- PossiblePossibly related (embedding) · 45%Transformer Geometry Observatory TGO-II: Representational Similarity Observatory →
- PossiblePossibly related (embedding) · 57%promptslab/Awesome-Prompt-Engineering →
Explains (incoming)
Related to (incoming)
Covers (incoming)
Related across the graph
paperThe State-Prediction Separation HypothesisnewsWhat performs the operations coordinated within each layer or head of a Transformer?paperLoop the Loopies!repolucidrains/x-transformersrepoquant-kitpaperTransformer Geometry Observatory TGO-II: Representational Similarity Observatoryrepopromptslab/Awesome-Prompt-EngineeringnewsI shrank a transformer until every number fitted on the screen and made the weights editable [R]paperGrokking in small transformerspaperGeneralization Analysis of Transformers in Distribution RegressionnewsTrain and run transformers directly on Apple's Neural Engine
