newsReddit r/MachineLearningTrust 52 · CommunityPublished 24d agoLive · 22d ago
SkewAdam: A tiered optimizer that cuts MoE state memory by 97% (fits a 6.7B MoE on a 40GB GPU) [R]
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 49%huggingface/optimum-intel →
- PossiblePossibly related (embedding) · 46%PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization →
