newsReddit r/MachineLearningTrust 52 · CommunityPublished 7d agoLive · 5d ago
I developed my own quantized LLM from scratch, trained on 30B tokens, deploys in 60 MB [R]
Reposting here after sharing this on [ r/MachineLearning ]( r/MachineLearning ) a few days ago, where it got a much better response than I expected (300+ upvotes, great questions, zero roasting) GitHub is at 35 stars now. So here it is. I trained a 250M parameter model from scratch on 30B tokens of fineweb. It’s quantized to under 2 bits so the whole deployment is 60 MB and it needs about
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 60%ModelEngine-Group/unified-cache-management →
- PossiblePossibly related (embedding) · 59%AetherAI3/Unlimited-Context-LLM →
- PossiblePossibly related (embedding) · 58%PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization →
- PossiblePossibly related (embedding) · 57%pythongiant/KVBoost →
- PossiblePossibly related (embedding) · 57%test5630352/llm-cache-optimize →
- PossiblePossibly related (embedding) · 49%jia-gao/leanctx →
- PossiblePossibly related (embedding) · 57%Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs →
Covers
Covers (incoming)
Related across the graph
paperQuantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMsrepotest5630352/llm-cache-optimizerepoModelEngine-Group/unified-cache-managementrepojia-gao/leanctxrepoAetherAI3/Unlimited-Context-LLMpaperPagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantizationrepopythongiant/KVBoost
