newsReddit r/LocalLLaMATrust 52 · CommunityPublished 8d agoLive · 8d ago
tried predicting which MoE experts get used next token to speed up cpu/gpu offload, got some real numbers, is this actually implementable or am i wasting my time (30tg/s -> 150-200tg/s)
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 49%ooples/token-optimizer-mcp →
- PossiblePossibly related (embedding) · 48%WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs →
- PossiblePossibly related (embedding) · 46%lucidrains/torch-einops-utils →
