Skip to main content
Angestrom home
SearchPapersModelsLive AIIntelligence
Search⌕⌘K
EnterprisePricingSign in

Stay Ahead in the AI Revolution

Weekly digest — EPI pulse, top intelligence, fresh lineage. Free, no account.

Follow Angestrom
Global source network
Synced every 5 minutes

Continuous sync from primary AI sources — indexed, enriched, and queryable in real time.

arXivHugging FaceGitHubOpenAIAnthropicDeepMindReutersBBC TechHacker NewsReddit MLVerified feedsFunding
ANGESTROM

The Intelligence Layer of Humanity. Everything AI. All in One Place.

Angestrom connects every piece of the AI ecosystem — data, models, research, companies, tools, and people.

info@angestrom.comwww.angestrom.comLucknow, Uttar Pradesh, India

Product

  • AI Search
  • AI Models
  • Research Papers
  • Companies
  • News & Events
  • GitHub Explorer
  • APIs & Tools
  • Datasets
  • Benchmarks
  • Model lifecycle
  • Funding graph
  • Contributors
  • AI Agents

Resources

  • Weekly digest
  • Documentation
  • Tutorials
  • Guides
  • News
  • Help / Start
  • Community

Company

  • About
  • Contact
  • Privacy Policy
  • Terms of Service
  • Acceptable Use

Enterprise

  • Pricing
  • Workspace
  • Contact Sales

Developer

  • Developer Hub
  • API docs
  • GitHub

Learn

  • Learning Academy
  • Roadmaps
  • Glossary
  • AI for Beginners

Popular Topics

Loading topics…
View All Topics →
© 2026 Angestrom Intelligence Private Limited. All rights reserved.
English
Theme
Angestrom home
SearchPapersModelsLive AIIntelligence
Search⌕⌘K
EnterprisePricingSign in
  1. Home
  2. /Repositories
  3. /lucidrains/x-transformers
Read original ↗
repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 25d ago

lucidrains/x-transformers

A concise but complete full-attention transformer with a set of promising experimental features from various papers

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 58%Build your first transformer from scratch →
  • PossiblePossibly related (embedding) · 48%Generalization Analysis of Transformers in Distribution Regression →
  • PossiblePossibly related (embedding) · 47%Morphing into Hybrid Attention Models →
  • FuzzySimilar title/name (fuzzy) · 87%Grokking in small transformers →

    “Fuzzy title match (0.94): “Grokking in small transformers” ≈ “lucidrains/x-transformers””

  • FuzzySimilar title/name (fuzzy) · 87%Kaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space Correlations →

    “Fuzzy title match (0.94): “Kaleido: Algorithm-Hardware Co-Design for Video Diffusion Tr” ≈ “lucidrains/x-transformers””

  • FuzzySimilar title/name (fuzzy) · 87%MxGPS: Multiplex Graph Transformers for a Power Grid Foundation Model →

    “Fuzzy title match (0.94): “MxGPS: Multiplex Graph Transformers for a Power Grid Foundat” ≈ “lucidrains/x-transformers””

  • FuzzySimilar title/name (fuzzy) · 87%FlexViT: A Flexible FPGA-based Accelerator for Edge Vision Transformers →

    “Fuzzy title match (0.94): “FlexViT: A Flexible FPGA-based Accelerator for Edge Vision T” ≈ “lucidrains/x-transformers””

  • FuzzySimilar title/name (fuzzy) · 87%Inhibited Self-Attention: Sharpening Focus in Vision Transformers →

    “Fuzzy title match (0.94): “Inhibited Self-Attention: Sharpening Focus in Vision Transfo” ≈ “lucidrains/x-transformers””

Related to

tutorialBuild your first transformer from scratch

Implements

paperGeneralization Analysis of Transformers in Distribution RegressionpaperMorphing into Hybrid Attention ModelspaperGrokking in small transformerspaperKaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space CorrelationspaperMxGPS: Multiplex Graph Transformers for a Power Grid Foundation ModelpaperFlexViT: A Flexible FPGA-based Accelerator for Edge Vision TransformerspaperInhibited Self-Attention: Sharpening Focus in Vision TransformerspaperAutomated Compliance Mapping in Cloud Security with Domain-Adapted Sentence TransformerspaperPAC-ACT: Post-training Actor-Critic for Action Chunking TransformerspaperReview Residuals: Update-Conditioned Residual Gating for TransformerspaperFoveation-Guided Dynamic Token Selection for Robust and Efficient Vision TransformerspaperFrom Expressivity to Sample Complexity: Narrow Teachers for Transformers via C-RASPpaperInvariant Learning Dynamics of Transformers in Inductive Reasoning TaskspaperPost-Training Pruning for Diffusion TransformerspaperUmm... With Transformers? Insights from Filled Pause Use across Four Slavic ParliamentspaperMobius Learning: Cyclic Depth Folding in TransformerspaperAppearance Pointers -- Multimodal Region Control of Diffusion TransformerspaperText Template Tokens Are Implicit Semantic Registers in Diffusion TransformerspaperELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training TransformerspaperKroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers

Related across the graph

paperReview Residuals: Update-Conditioned Residual Gating for TransformerspaperAutomated Compliance Mapping in Cloud Security with Domain-Adapted Sentence TransformerspaperInhibited Self-Attention: Sharpening Focus in Vision TransformerspaperAppearance Pointers -- Multimodal Region Control of Diffusion TransformerspaperPAC-ACT: Post-training Actor-Critic for Action Chunking TransformerspaperFrom Expressivity to Sample Complexity: Narrow Teachers for Transformers via C-RASPpaperFoveation-Guided Dynamic Token Selection for Robust and Efficient Vision TransformerspaperFlexViT: A Flexible FPGA-based Accelerator for Edge Vision TransformerspaperELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training TransformerspaperText Template Tokens Are Implicit Semantic Registers in Diffusion TransformerspaperKaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space CorrelationspaperKroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion TransformerspaperMxGPS: Multiplex Graph Transformers for a Power Grid Foundation ModelpaperMorphing into Hybrid Attention ModelspaperInvariant Learning Dynamics of Transformers in Inductive Reasoning TaskspaperMobius Learning: Cyclic Depth Folding in TransformerstutorialBuild your first transformer from scratchpaperUmm... With Transformers? Insights from Filled Pause Use across Four Slavic ParliamentspaperGrokking in small transformerspaperGeneralization Analysis of Transformers in Distribution RegressionpaperPost-Training Pruning for Diffusion Transformers
Knowledge path·PReview Residuals: Update-Conditioned Residual Gating for Transformers→PAutomated Compliance Mapping in Cloud Security with Domain-Adapted Sentence Transformers→PInhibited Self-Attention: Sharpening Focus in Vision Transformers→Rlucidrains/x-transformers

Topics

artificial-intelligenceattention-mechanismdeep-learningtransformers

Explore

Search similar →Knowledge graph →All repos →Full intelligence feed →
Graph trust82Primary
Graph score5922