newsReddit r/MachineLearningTrust 52 · CommunityPublished 1mo agoLive · 1mo ago
H64LM: A 249M-parameter Mixture-of-Experts Transformer built from scratch in PyTorch [P]
Hi everyone, I built H64LM, a research project to better understand modern LLMs by implementing one from scratch in PyTorch. Instead of relying on high-level training frameworks, I implemented the core components myself attention, MoE routing, normalization, and the training loop. Features 249M-parameter Transformer Grouped Query Attention (GQA) Sparse Mixture-of-Experts (8 experts, Top-2 routi
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%luziyao1995/vllm →
- PossiblePossibly related (embedding) · 52%alibaba/rtp-llm →
- PossiblePossibly related (embedding) · 49%peremartra/Rearchitecting-LLMs →
- PossiblePossibly related (embedding) · 49%quant-kit →
- PossiblePossibly related (embedding) · 48%Understanding Large Language Models →
- PossiblePossibly related (embedding) · 49%ludwig-ai/ludwig →
- PossiblePossibly related (embedding) · 48%lucidrains/torch-einops-utils →
- PossiblePossibly related (embedding) · 47%kyegomez/BitNet →
Covers
Covers (incoming)
repolucidrains/torch-einops-utilsrepokyegomez/BitNetrepoluojieLLMaaS/haxivrepooptuna/optunareporasbt/reasoning-from-scratchpaperSuper Weights in LLMs and the Failure of Selective TrainingpaperIt Takes a MAESTRO To Prune Bad Expertsrepomodelscope/mcore-bridgerepolibxsmm/libxsmmrepoAarambhDevHub/aarambh-airepoOptim-Agent/optim-agentpaperLoop the Loopies!paperManifold-Constrained Hyper-Connections for Parameter-Efficient Finetuningreposkorch-dev/skorchrepomicrosoft/TutelpaperELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training Transformersrepohiyouga/EasyR1repoAarambhDevHub/aarambh-studiopaperDeaMoE: Efficient MoE Structure for Fast Small-Batch Decoding
Related across the graph
repolucidrains/torch-einops-utilspaperELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training TransformerspaperLoop the Loopies!repoalibaba/rtp-llmrepoludwig-ai/ludwigpaperIt Takes a MAESTRO To Prune Bad Expertsrepokyegomez/BitNetrepoquant-kitrepooptuna/optunarepomodelscope/mcore-bridgepaperManifold-Constrained Hyper-Connections for Parameter-Efficient FinetuningrepoOptim-Agent/optim-agentrepoAarambhDevHub/aarambh-airepolibxsmm/libxsmmrepoluziyao1995/vllmpaperSuper Weights in LLMs and the Failure of Selective TrainingrepoluojieLLMaaS/haxivpaperUnderstanding Large Language Modelsrepoperemartra/Rearchitecting-LLMsrepoAarambhDevHub/aarambh-studioreporasbt/reasoning-from-scratchrepomicrosoft/Tutelreposkorch-dev/skorchpaperDeaMoE: Efficient MoE Structure for Fast Small-Batch Decodingrepohiyouga/EasyR1
